fix: decode extractor responses with declared charset - #87
Merged
Merged
Conversation
Member
|
Thx a lot for ur contribution, merging to dev :) ! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Searching for anything non-ASCII turned the suggestions into � (U+FFFD) chars. Googles suggestions endpoint answers in a locale dependent charset (
ISO-8859-1,Shift_JISorwindows-1251), butOkHttpExtractorResponseMapperdecoded every body asUTF-8.The mapper now decodes with the charset from the response
Content-Type, falling back toUTF-8when nothing is declared, so all other traffic is unchanged.The mapper sits in front of all extractor traffic, but all the other endpoints already answers in
UTF-8afaik, so really only the suggestions endpoint is effected here.Added regression tests in
OkHttpExtractorResponseMapperTest.To test the upstream endpoint:
Ë(%C3%8B)ISO-8859-10xCBbyteThe charset also seems to follow the requesters locale:
First returns
Shift_JISand second returnswindows-1251. Just here to demonstrate locale dependency. We never send one though, soISO-8859-1is all we ever see.