a ‚oeÛTã @s´ddlZddlmZddlmZmZmZmZmZddl m Z m Z m Z m Z ddlmZmZmZmZddlmZddlmZmZdd lmZmZmZmZmZmZmZe  d ¡Z!e "¡Z#e# $e %d ¡¡dee&e'fe(e(e)eee*eee*e+e+e)e+edœ dd„Z,dee(e(e)eee*eee*e+e+e)e+edœ dd„Z-d ee*e&efe(e(e)eee*eee*e+e+e)e+edœ dd„Z.d!eee*ee&fe(e(e)eee*eee*e+e+e)e+e+dœ dd„Z/dS)"éN)ÚPathLike)ÚBinaryIOÚListÚOptionalÚSetÚUnioné)Úcoherence_ratioÚencoding_languagesÚmb_encoding_languagesÚmerge_coherence_ratios)ÚIANA_SUPPORTEDÚTOO_BIG_SEQUENCEÚTOO_SMALL_SEQUENCEÚTRACE)Ú mess_ratio)Ú CharsetMatchÚCharsetMatches)Úany_specified_encodingÚcut_sequence_chunksÚ iana_nameÚidentify_sig_or_bomÚ is_cp_similarÚis_multi_byte_encodingÚshould_strip_sig_or_bomZcharset_normalizerz)%(asctime)s | %(levelname)s | %(message)sééçš™™™™™É?TFçš™™™™™¹?) Ú sequencesÚstepsÚ chunk_sizeÚ thresholdÚ cp_isolationÚ cp_exclusionÚpreemptive_behaviourÚexplainÚlanguage_thresholdÚenable_fallbackÚreturnc / Cs t|ttfƒs td t|ƒ¡ƒ‚|r>tj} t t ¡t  t ¡t |ƒ} | dkrŽt  d¡|rvt t ¡t  | prtj¡tt|dddgdƒgƒS|durºt t d d  |¡¡d d „|Dƒ}ng}|durêt t d d  |¡¡dd „|Dƒ}ng}| ||k�rt t d||| ¡d}| }|dk�r:| ||k�r:t| |ƒ}t |ƒtk} t |ƒtk} | �rlt t d | ¡¡n| �r„t t d | ¡¡g}|�r–t|ƒnd}|du�r¼| |¡t t d|¡tƒ}g}g}d}d}d}tƒ}t|ƒ\}}|du�r| |¡t t dt |ƒ|¡| d¡d|v�r.| d¡|tD�]Þ}|�rP||v�rP�q6|�rd||v�rd�q6||v�rr�q6| |¡d}||k}|�o”t|ƒ}|dv�r¸|�s¸t t d|¡�q6|dv�rÚ|�sÚt t d|¡�q6z t|ƒ}Wn,t t!f�yt t d|¡Y�q6Yn0zr| �r^|du�r^t"|du�rB|dtdƒ…n|t |ƒtdƒ…|d�n&t"|du�rn|n|t |ƒd…|d�}Wnbt#t$f�yè}zDt|t$ƒ�s¼t t d|t"|ƒ¡| |¡WYd}~�q6WYd}~n d}~00d}|D]} t%|| ƒ�ròd}�q�qò|�r*t t d|| ¡�q6t&|�s6dnt |ƒ| t| |ƒƒ}!|�of|du�oft |ƒ| k}"|"�r|t t d |¡tt |!ƒd!ƒ}#t'|#d"ƒ}#d}$d}%g}&g}'zšt(|||!||||||ƒ D]|}(|& |(¡|' t)|(||du�oþdt |ƒk�oúd"knƒ¡|'d#|k�r|$d7}$|$|#k�s4|�rÀ|du�rÀ�q>�qÀWnBt#�y‚}z(t t d$|t"|ƒ¡|#}$d}%WYd}~n d}~00|%�s| �r|�sz|td%ƒd…j*|d&d'�WnRt#�y}z8t t d(|t"|ƒ¡| |¡WYd}~�q6WYd}~n d}~00|'�rt+|'ƒt |'ƒnd})|)|k�s6|$|#k�r´| |¡t t d)||$t,|)d*d+d,�¡| �r6|dd|fv�r6|%�s6t|||dg|ƒ}*||k�rœ|*}n|dk�r¬|*}n|*}�q6t t d-|t,|)d*d+d,�¡|�sàt-|ƒ}+nt.|ƒ}+|+�rt t d. |t"|+ƒ¡¡g},|dk�rF|&D],}(t/|(||+�r2d/ |+¡ndƒ}-|, |-¡�qt0|,ƒ}.|.�rht t d0 |.|¡¡| t|||)||.|ƒ¡||ddfv�rÒ|)d1k�rÒt  d2|¡|�rÀt t ¡t  | ¡t||gƒS||k�r6t  d3|¡|�rt t ¡t  | ¡t||gƒS�q6t |ƒdk�rÈ|�s8|�s8|�rDt t d4¡|�rdt  d5|j1¡| |¡nd|�rt|du�s˜|�rŽ|�rŽ|j2|j2k�s˜|du�r®t  d6¡| |¡n|�rÈt  d7¡| |¡|�rìt  d8| 3¡j1t |ƒd¡n t  d9¡|� rt t ¡t  | ¡|S):af Given a raw bytes sequence, return the best possibles charset usable to render str objects. If there is no results, it is a strong indicator that the source is binary/not text. By default, the process will extract 5 blocks of 512o each to assess the mess and coherence of a given sequence. And will give up a particular code page after 20% of measured mess. Those criteria are customizable at will. The preemptive behavior DOES NOT replace the traditional detection workflow, it prioritize a particular code page but never take it for granted. Can improve the performance. You may want to focus your attention to some code page or/and not others, use cp_isolation and cp_exclusion for that purpose. This function will strip the SIG in the payload/sequence every time except on UTF-16, UTF-32. By default the library does not setup any handler other than the NullHandler, if you choose to set the 'explain' toggle to True it will alter the logger configuration to add a StreamHandler that is suitable for debugging. Custom logging format and handler can be set manually. z4Expected object of type bytes or bytearray, got: {0}rz[ózfrom_bytes..zacp_exclusion is set. use this flag for debugging purpose. limited list of encoding excluded : %s.cSsg|]}t|dƒ‘qSr,r-r.r0r0r1r2fr3z^override steps (%i) and chunk_size (%i) as content does not fit (%i byte(s) given) parameters.rz>Trying to detect encoding from a tiny portion of ({}) byte(s).zIUsing lazy str decoding because the payload is quite large, ({}) byte(s).z@Detected declarative mark in sequence. Priority +1 given for %s.zIDetected a SIG or BOM mark on first %i byte(s). Priority +1 given for %s.Úascii>Úutf_16Úutf_32z\Encoding %s won't be tested as-is because it require a BOM. Will try some sub-encoder LE/BE.>Úutf_7zREncoding %s won't be tested as-is because detection is unreliable without BOM/SIG.z2Encoding %s does not provide an IncrementalDecoderg€„A)Úencodingz9Code page %s does not fit given bytes sequence at ALL. %sTzW%s is deemed too similar to code page %s and was consider unsuited already. Continuing!zpCode page %s is a multi byte encoding table and it appear that at least one character was encoded using n-bytes.éééÿÿÿÿzaLazyStr Loading: After MD chunk decode, code page %s does not fit given bytes sequence at ALL. %sgjè@Ústrict)Úerrorsz^LazyStr Loading: After final lookup, code page %s does not fit given bytes sequence at ALL. %szc%s was excluded because of initial chaos probing. Gave up %i time(s). Computed mean chaos is %f %%.édé)Úndigitsz=%s passed initial chaos probing. Mean measured chaos is %f %%z&{} should target any language(s) of {}ú,z We detected language {} using {}rz.Encoding detection: %s is most likely the one.zoEncoding detection: %s is most likely the one as we detected a BOM or SIG within the beginning of the sequence.zONothing got out of the detection process. Using ASCII/UTF-8/Specified fallback.z7Encoding detection: %s will be used as a fallback matchz:Encoding detection: utf_8 will be used as a fallback matchz:Encoding detection: ascii will be used as a fallback matchz]Encoding detection: Found %s as plausible (best-candidate) for content. With %i alternatives.z=Encoding detection: Unable to determine any suitable charset.)4Ú isinstanceÚ bytearrayÚbytesÚ TypeErrorÚformatÚtypeÚloggerÚlevelZ addHandlerÚexplain_handlerZsetLevelrÚlenÚdebugZ removeHandlerÚloggingZWARNINGrrÚlogÚjoinÚintrrrÚappendÚsetrr ÚaddrrÚModuleNotFoundErrorÚ ImportErrorÚstrÚUnicodeDecodeErrorÚ LookupErrorrÚrangeÚmaxrrÚdecodeÚsumÚroundr r r r r8Z fingerprintZbest)/rr r!r"r#r$r%r&r'r(Zprevious_logger_levelÚlengthZis_too_small_sequenceZis_too_large_sequenceZprioritized_encodingsZspecified_encodingZtestedZtested_but_hard_failureZtested_but_soft_failureZfallback_asciiZ fallback_u8Zfallback_specifiedÚresultsZ sig_encodingZ sig_payloadZ encoding_ianaZdecoded_payloadZbom_or_sig_availableZstrip_sig_or_bomZis_multi_byte_decoderÚeZsimilar_soft_failure_testZencoding_soft_failedZr_Zmulti_byte_bonusZmax_chunk_gave_upZearly_stop_countZlazy_str_hard_failureZ md_chunksZ md_ratiosÚchunkZmean_mess_ratioZfallback_entryZtarget_languagesZ cd_ratiosZchunk_languagesZcd_ratios_mergedr0r0r1Ú from_bytes!sâÿÿ    üüû   ÿþÿþÿ  ý   ü     ÿýý ý ÿüÿü  ü $  ü ýÿ ýü ÷ &ýÿ ÿÿÿ üÿþýü $ ú ÿ þý ÿ  ü ÿþ ýÿþúÿ ÿþÿ   ý  þþ ÿÿýü ûù     ý   rb) Úfpr r!r"r#r$r%r&r'r(r)c Cst| ¡||||||||| ƒ S)z† Same thing than the function from_bytes but using a file pointer that is already ready. Will not close the file pointer. )rbÚread) rcr r!r"r#r$r%r&r'r(r0r0r1Úfrom_fpösöre) Úpathr r!r"r#r$r%r&r'r(r)c CsHt|dƒ�*} t| ||||||||| ƒ WdƒS1s:0YdS)z• Same thing than the function from_bytes but with one extra step. Opening and reading given file path in binary mode. Can raise IOError. ÚrbN)Úopenre) rfr r!r"r#r$r%r&r'r(rcr0r0r1Ú from_paths öri) Úfp_or_path_or_payloadr r!r"r#r$r%r&r'r(r)c Cszt|ttfƒr,t|||||||||| d� } nHt|ttfƒrXt|||||||||| d� } nt|||||||||| d� } | S)a) Detect if the given input (file, bytes, or path) points to a binary file. aka. not a string. Based on the same main heuristic algorithms and default kwargs at the sole exception that fallbacks match are disabled to be stricter around ASCII-compatible but unlikely to be a string. ) r r!r"r#r$r%r&r'r()rBrVrrirDrCrbre) rjr r!r"r#r$r%r&r'r(Zguessesr0r0r1Ú is_binary3sXö þþö ö rk) rrrNNTFrT) rrrNNTFrT) rrrNNTFrT) rrrNNTFrF)0rMÚosrÚtypingrrrrrZcdr r r r Zconstantr rrrZmdrZmodelsrrZutilsrrrrrrrZ getLoggerrHZ StreamHandlerrJZ setFormatterZ FormatterrDrCrPÚfloatrVÚboolrbrerirkr0r0r0r1ÚsÎ  $ ÿö   õ Zö  õ ö   õ !ö  õ