Skip to content

Latest commit

 

History

History
683 lines (544 loc) · 27.2 KB

File metadata and controls

683 lines (544 loc) · 27.2 KB

Decoder API Reference

This document follows the current public headers under include/. decoder.h is the umbrella header. It re-exports the low-level modules and adds a separate UTF-8-first pipeline API.

Header Map

  • decoder.h: umbrella include + high-level UTF-8 pipeline API
  • core.h: library lifecycle and diagnostics
  • types.h: public enums, structs, version constants
  • properties.h: Unicode character properties
  • encoding.h: UTF-8 / UTF-16 / UTF-32 validation and conversion
  • case.h: case mapping and case folding
  • normalize.h: Unicode normalization
  • segment.h: grapheme / word / sentence segmentation
  • script.h: script and block detection
  • emoji.h: emoji properties and sequence parsing
  • security.h: Unicode security and confusables
  • parallel.h: multi-threaded helpers

Common Conventions

  • Low-level Unicode algorithms mostly operate on UTF-32 code point arrays.
  • High-level functions in decoder.h operate on UTF-8 byte buffers.
  • Functions that return int use decoder_status_t values from types.h.
  • Output buffers never write past capacity. When a function documents DECODER_ERROR_BUFFER_TOO_SMALL, inspect its output length parameter for the number of units written or required.
  • Some primitive APIs intentionally return size_t instead of a status code: this is used for counts, direct navigation, or fixed-size decomposition / mapping helpers.

Core Types and Constants (types.h)

Version Constants

  • DECODER_VERSION_MAJOR
  • DECODER_VERSION_MINOR
  • DECODER_VERSION_PATCH
  • DECODER_VERSION_STRING
  • DECODER_STANDARD_VERSION
  • DECODER_VERSION_FULL
  • DECODER_ABI_VERSION
  • DECODER_VERSION_CHECK(major, minor, patch)
  • DECODER_FEATURE_SIMD
  • DECODER_FEATURE_PARALLEL
  • DECODER_BLOCK_NO_BLOCK

Status Codes

decoder_status_t:

  • DECODER_SUCCESS
  • DECODER_ERROR_INVALID_INPUT
  • DECODER_ERROR_BUFFER_TOO_SMALL
  • DECODER_ERROR_INVALID_UTF8
  • DECODER_ERROR_INVALID_UTF16
  • DECODER_ERROR_INVALID_CODEPOINT
  • DECODER_ERROR_OUT_OF_MEMORY
  • DECODER_ERROR_NOT_IMPLEMENTED
  • DECODER_ERROR_IO
  • DECODER_ERROR_INVALID_ARGUMENT
  • DECODER_ERROR_OVERFLOW
  • DECODER_ERROR_NOT_FOUND

Common Public Enums

  • decoder_category_t
  • decoder_normalization_form_t
  • decoder_quick_check_t
  • decoder_script_t
  • decoder_script_family_t
  • decoder_grapheme_break_t
  • decoder_word_break_t
  • decoder_sentence_break_t
  • decoder_confusable_type_t
  • decoder_block_t
  • decoder_emoji_skin_tone_t
  • decoder_emoji_sequence_type_t
  • decoder_indic_syllabic_category_t
  • decoder_indic_positional_category_t
  • decoder_script_resolution_mode_t
  • decoder_script_priority_t

Common Public Structs

Locale and analysis:

  • decoder_locale_t
    • language
    • country
    • variant
    • encoding
  • decoder_script_analysis_t
    • scripts
    • count
    • has_common
    • has_inherited
    • is_suspicious
  • decoder_script_run_t
    • start
    • end
    • script
  • decoder_script_range_t
    • start
    • end
    • script
    • priority
    • ucd_authoritative
    • source
  • decoder_script_conflict_result_t
    • start
    • end
    • resolved_script
  • decoder_script_unification_t
    • ranges
    • range_count
    • range_capacity
    • mode
    • strict_mode
  • DECODER_SCRIPT_LABEL_MAX (label/source string capacity)
  • decoder_security_options_t
  • decoder_homoglyph_result_t
    • homoglyphs
    • count

Segmentation and emoji:

  • decoder_segment_t
    • start
    • end
  • decoder_grapheme_iter_t
  • decoder_word_iter_t
  • decoder_sentence_iter_t
  • decoder_emoji_sequence_info_t
    • type
    • length
    • base_emoji
    • skin_tone
    • has_variation_selector
    • has_zwj
  • decoder_boundary_stats_t
    • total_chars
    • grapheme_count
    • word_count
    • sentence_count
    • whitespace_count
    • punctuation_count

Core Lifecycle (core.h)

  • void decoder_init(void)
  • void decoder_cleanup(void)
  • const char *decoder_get_unicode_version(void)
  • const char *decoder_error_message(decoder_status_t status)

Call decoder_init() before using the library and decoder_cleanup() when the process no longer needs it. Both are intended to be safe to call repeatedly.

Character Properties (properties.h)

Validity and Classification

  • bool decoder_is_valid(uint32_t cp)
  • bool decoder_is_assigned(uint32_t cp)
  • bool decoder_is_private_use(uint32_t cp)
  • bool decoder_is_surrogate(uint32_t cp)
  • bool decoder_is_noncharacter(uint32_t cp)
  • decoder_category_t decoder_get_category(uint32_t cp)

Derived Predicates

  • bool decoder_is_letter(uint32_t cp)
  • bool decoder_is_uppercase(uint32_t cp)
  • bool decoder_is_lowercase(uint32_t cp)
  • bool decoder_is_titlecase(uint32_t cp)
  • bool decoder_is_digit(uint32_t cp)
  • bool decoder_is_number(uint32_t cp)
  • bool decoder_is_punctuation(uint32_t cp)
  • bool decoder_is_symbol(uint32_t cp)
  • bool decoder_is_mark(uint32_t cp)
  • bool decoder_is_separator(uint32_t cp)
  • bool decoder_is_control(uint32_t cp)
  • bool decoder_is_format(uint32_t cp)
  • bool decoder_is_whitespace(uint32_t cp)
  • bool decoder_is_space(uint32_t cp)
  • bool decoder_is_alphanumeric(uint32_t cp)
  • bool decoder_is_alphabetic(uint32_t cp)
  • bool decoder_is_numeric(uint32_t cp)

Numeric Values and Metadata

  • double decoder_get_numeric_value(uint32_t cp)
  • int decoder_get_digit_value(uint32_t cp)
  • int decoder_get_name(uint32_t cp, char *buffer, size_t capacity)
  • uint32_t decoder_from_name(const char *name)
  • const char *decoder_get_age(uint32_t cp)
  • bool decoder_is_in_version(uint32_t cp, int major, int minor)

decoder_get_name() returns DECODER_SUCCESS, DECODER_ERROR_BUFFER_TOO_SMALL, or DECODER_ERROR_NOT_FOUND. Coverage includes explicit names from UnicodeData.txt, CJK Unified and Tangut algorithmic ranges, and Hangul syllables computed from L/V/T jamo. decoder_from_name() is case-insensitive, recognises CJK UNIFIED IDEOGRAPH-{HEX} and TANGUT IDEOGRAPH-{HEX} prefixes, honours NameAliases.txt, and returns 0xFFFFFFFF on miss. decoder_get_age() returns NULL for unassigned code points.

Indic Properties (UAX #44)

  • decoder_indic_syllabic_category_t decoder_get_indic_syllabic_category(uint32_t cp)
  • decoder_indic_positional_category_t decoder_get_indic_positional_category(uint32_t cp)
  • const char *decoder_get_indic_syllabic_category_name(decoder_indic_syllabic_category_t category)
  • const char *decoder_get_indic_positional_category_name(decoder_indic_positional_category_t category)

Code points outside Indic scripts report DECODER_INDIC_SYLLABIC_OTHER / DECODER_INDIC_POSITIONAL_NA. The name accessors return static strings such as "Vowel_Independent" or "Top_And_Bottom".

Batch Classification

  • decoder_charclass_t
    • DECODER_CHARCLASS_LETTER
    • DECODER_CHARCLASS_DIGIT
    • DECODER_CHARCLASS_WHITESPACE
    • DECODER_CHARCLASS_PUNCTUATION
    • DECODER_CHARCLASS_SYMBOL
    • DECODER_CHARCLASS_NEWLINE
    • DECODER_CHARCLASS_OTHER
  • int decoder_classify_codepoints(const uint8_t *utf8, size_t utf8_len, uint8_t *classes_out, size_t classes_cap, size_t *codepoint_count)

Encoding Conversion (encoding.h)

UTF-8

  • bool decoder_is_valid_utf8(const uint8_t *str, size_t len)
  • size_t decoder_utf8_length(const uint8_t *str, size_t byte_len)
  • size_t decoder_utf8_char_count(const uint8_t *str, size_t byte_len)
  • int decoder_utf8_to_utf32(const uint8_t *src, size_t src_len, uint32_t *dst, size_t dst_capacity, size_t *chars_written)
  • int decoder_utf32_to_utf8(const uint32_t *src, size_t src_len, uint8_t *dst, size_t dst_capacity, size_t *bytes_written)
  • int decoder_utf8_to_utf16(const uint8_t *src, size_t src_len, uint16_t *dst, size_t dst_capacity, size_t *units_written)
  • int decoder_utf16_to_utf8(const uint16_t *src, size_t src_len, uint8_t *dst, size_t dst_capacity, size_t *bytes_written)

UTF-16 / UTF-32

  • bool decoder_is_valid_utf16(const uint16_t *str, size_t len)
  • int decoder_utf16_to_utf32(const uint16_t *src, size_t src_len, uint32_t *dst, size_t dst_capacity, size_t *chars_written)
  • int decoder_utf32_to_utf16(const uint32_t *src, size_t src_len, uint16_t *dst, size_t dst_capacity, size_t *units_written)
  • bool decoder_is_valid_utf32(const uint32_t *str, size_t len)

Case Mapping (case.h)

Locale-sensitive support is intentionally narrow:

  • Turkic dotted / dotless I rules are supported in locale-aware APIs.
  • Greek final sigma is handled in lowercase and titlecase paths.
  • Other locale tags fall back to default Unicode mappings.

Single Code Point

  • uint32_t decoder_to_upper(uint32_t cp)
  • uint32_t decoder_to_lower(uint32_t cp)
  • uint32_t decoder_to_title(uint32_t cp)
  • uint32_t decoder_case_fold(uint32_t cp)

Full 1:N Mapping

  • size_t decoder_to_upper_full(uint32_t cp, uint32_t *out, size_t capacity)
  • size_t decoder_to_lower_full(uint32_t cp, uint32_t *out, size_t capacity)
  • size_t decoder_to_title_full(uint32_t cp, uint32_t *out, size_t capacity)
  • size_t decoder_case_fold_full(uint32_t cp, uint32_t *out, size_t capacity)

These functions return 0 when the provided output buffer cannot hold the mapping.

UTF-32 Strings

  • int decoder_string_to_upper(const uint32_t *src, size_t src_len, uint32_t *dst, size_t dst_capacity, size_t *dst_len)
  • int decoder_string_to_lower(const uint32_t *src, size_t src_len, uint32_t *dst, size_t dst_capacity, size_t *dst_len)
  • int decoder_string_to_title(const uint32_t *src, size_t src_len, uint32_t *dst, size_t dst_capacity, size_t *dst_len)
  • int decoder_string_case_fold(const uint32_t *src, size_t src_len, uint32_t *dst, size_t dst_capacity, size_t *dst_len)
  • int decoder_string_to_upper_locale(const uint32_t *src, size_t src_len, const decoder_locale_t *locale, uint32_t *dst, size_t dst_capacity, size_t *dst_len)
  • int decoder_string_to_lower_locale(const uint32_t *src, size_t src_len, const decoder_locale_t *locale, uint32_t *dst, size_t dst_capacity, size_t *dst_len)
  • int decoder_string_case_fold_locale(const uint32_t *src, size_t src_len, const decoder_locale_t *locale, uint32_t *dst, size_t dst_capacity, size_t *dst_len)

UTF-8 Strings

  • int decoder_utf8_to_upper(const uint8_t *src, size_t src_len, uint8_t *dst, size_t dst_capacity, size_t *dst_len)
  • int decoder_utf8_to_lower(const uint8_t *src, size_t src_len, uint8_t *dst, size_t dst_capacity, size_t *dst_len)
  • int decoder_utf8_case_fold(const uint8_t *src, size_t src_len, uint8_t *dst, size_t dst_capacity, size_t *dst_len)

Predicates and Comparison

  • bool decoder_is_case_ignorable(uint32_t cp)
  • bool decoder_is_cased(uint32_t cp)
  • int decoder_case_compare(const uint32_t *s1, size_t len1, const uint32_t *s2, size_t len2)
  • bool decoder_changes_when_uppercased(uint32_t cp)
  • bool decoder_changes_when_lowercased(uint32_t cp)
  • bool decoder_changes_when_titlecased(uint32_t cp)
  • bool decoder_changes_when_casefolded(uint32_t cp)
  • bool decoder_case_equal(const uint32_t *s1, size_t len1, const uint32_t *s2, size_t len2)

Normalization (normalize.h)

UTF-32

  • int decoder_normalize(const uint32_t *src, size_t src_len, decoder_normalization_form_t form, uint32_t *dst, size_t dst_capacity, size_t *dst_len)
  • bool decoder_is_normalized(const uint32_t *str, size_t len, decoder_normalization_form_t form)
  • decoder_quick_check_t decoder_quick_check(const uint32_t *str, size_t len, decoder_normalization_form_t form)
  • int decoder_normalize_compare(const uint32_t *s1, size_t len1, const uint32_t *s2, size_t len2, decoder_normalization_form_t form)

decoder_normalize() writes the required output length to dst_len when it returns DECODER_ERROR_BUFFER_TOO_SMALL.

UTF-8 Convenience Wrappers

  • int decoder_normalize_utf8(const uint8_t *src, size_t src_len, decoder_normalization_form_t form, uint8_t *dst, size_t dst_capacity, size_t *dst_len)
  • bool decoder_is_normalized_utf8(const uint8_t *str, size_t len, decoder_normalization_form_t form)

Low-Level Primitives

  • uint8_t decoder_get_combining_class(uint32_t cp)
  • bool decoder_is_combining(uint32_t cp)
  • bool decoder_can_compose(uint32_t a, uint32_t b)
  • uint32_t decoder_compose(uint32_t a, uint32_t b)
  • size_t decoder_decompose(uint32_t cp, uint32_t *out, size_t capacity)
  • size_t decoder_decompose_compat(uint32_t cp, uint32_t *out, size_t capacity)

Text Segmentation (segment.h)

All segmentation APIs work on UTF-32 code point arrays and implement UAX #29.

Stateless Navigation

Grapheme clusters:

  • size_t decoder_next_grapheme(const uint32_t *str, size_t len, size_t pos)
  • size_t decoder_prev_grapheme(const uint32_t *str, size_t len, size_t pos)
  • size_t decoder_count_graphemes(const uint32_t *str, size_t len)
  • bool decoder_is_grapheme_boundary(const uint32_t *str, size_t len, size_t pos)
  • int decoder_find_grapheme_boundaries(const uint32_t *text, size_t len, size_t *boundaries, size_t boundaries_cap, size_t *boundaries_len)

Words:

  • size_t decoder_next_word(const uint32_t *str, size_t len, size_t pos)
  • size_t decoder_prev_word(const uint32_t *str, size_t len, size_t pos)
  • size_t decoder_count_words(const uint32_t *str, size_t len)
  • bool decoder_is_word_boundary(const uint32_t *str, size_t len, size_t pos)
  • int decoder_find_word_boundaries(const uint32_t *text, size_t len, size_t *boundaries, size_t boundaries_cap, size_t *boundaries_len)

Sentences:

  • size_t decoder_next_sentence(const uint32_t *str, size_t len, size_t pos)
  • size_t decoder_prev_sentence(const uint32_t *str, size_t len, size_t pos)
  • size_t decoder_count_sentences(const uint32_t *str, size_t len)
  • bool decoder_is_sentence_boundary(const uint32_t *str, size_t len, size_t pos)
  • int decoder_find_sentence_boundaries(const uint32_t *text, size_t len, size_t *boundaries, size_t boundaries_cap, size_t *boundaries_len)

The boundary-finding APIs emit positions starting at 0 and ending at len when the provided buffer is large enough.

Iterators

  • void decoder_grapheme_iter_init(decoder_grapheme_iter_t *it, const uint32_t *text, size_t len)
  • bool decoder_grapheme_iter_next(decoder_grapheme_iter_t *it, decoder_segment_t *seg)
  • void decoder_grapheme_iter_reset(decoder_grapheme_iter_t *it)
  • void decoder_word_iter_init(decoder_word_iter_t *it, const uint32_t *text, size_t len)
  • bool decoder_word_iter_next(decoder_word_iter_t *it, decoder_segment_t *seg)
  • void decoder_word_iter_reset(decoder_word_iter_t *it)
  • void decoder_sentence_iter_init(decoder_sentence_iter_t *it, const uint32_t *text, size_t len)
  • bool decoder_sentence_iter_next(decoder_sentence_iter_t *it, decoder_segment_t *seg)
  • void decoder_sentence_iter_reset(decoder_sentence_iter_t *it)

Statistics

  • int decoder_get_boundary_stats(const uint32_t *str, size_t len, decoder_boundary_stats_t *stats)

Script and Block Detection (script.h)

Per-Code-Point Queries

  • decoder_script_t decoder_get_script(uint32_t cp)
  • const char *decoder_get_script_name(decoder_script_t script)
  • const char *decoder_get_script_description(decoder_script_t script)
  • bool decoder_is_valid_script(decoder_script_t script)
  • decoder_script_family_t decoder_get_script_family(decoder_script_t script)
  • const char *decoder_get_script_family_name(uint32_t cp)
  • const char *decoder_get_block_name(uint32_t cp)
  • decoder_block_t decoder_get_block(uint32_t cp)
  • bool decoder_is_in_block(uint32_t cp, decoder_block_t block)
  • int decoder_get_block_range(decoder_block_t block, uint32_t *start, uint32_t *end)

The family taxonomy is informal: Unicode does not standardise script families. Choices follow cultural-region heuristics (e.g. Cherokee → American, N'Ko → African). See the decoder_script_family_t doc-comment in types.h for the full mapping rationale.

Script-Specific Checks

  • bool decoder_is_latin(uint32_t cp)
  • bool decoder_is_cyrillic(uint32_t cp)
  • bool decoder_is_greek(uint32_t cp)
  • bool decoder_is_arabic(uint32_t cp)
  • bool decoder_is_hebrew(uint32_t cp)
  • bool decoder_is_devanagari(uint32_t cp)
  • bool decoder_is_thai(uint32_t cp)
  • bool decoder_is_cjk(uint32_t cp)
  • bool decoder_is_hangul(uint32_t cp)
  • bool decoder_is_hiragana(uint32_t cp)
  • bool decoder_is_katakana(uint32_t cp)

String-Level Analysis

  • int decoder_analyze_scripts(const uint32_t *str, size_t len, decoder_script_analysis_t *analysis)
  • void decoder_script_analysis_free(decoder_script_analysis_t *analysis)
  • int decoder_detect_script_runs(const uint32_t *text, size_t length, decoder_script_run_t **runs, size_t *count)
  • decoder_script_t decoder_get_primary_script(const uint32_t *str, size_t len)
  • size_t decoder_count_scripts(const uint32_t *str, size_t len)
  • bool decoder_is_mixed_script(const uint32_t *str, size_t len)

Script Unification Engine

A mutable, in-memory range manager separate from the static UCD lookup, intended for tooling that combines UCD data with custom or legacy mappings and arbitrates overlaps by priority.

  • bool decoder_script_unification_init(decoder_script_unification_t *engine, decoder_script_resolution_mode_t mode, bool strict_mode)
  • void decoder_script_unification_cleanup(decoder_script_unification_t *engine)
  • bool decoder_script_unification_add_range(decoder_script_unification_t *engine, uint32_t start, uint32_t end, const char *script, decoder_script_priority_t priority, bool ucd_authoritative, const char *source)
  • bool decoder_script_unification_resolve_conflicts(decoder_script_unification_t *engine, decoder_script_conflict_result_t *conflicts, size_t max_conflicts, size_t *conflict_count)
  • const char *decoder_script_unification_get_script(const decoder_script_unification_t *engine, uint32_t cp)
  • bool decoder_script_unification_find_gaps(const decoder_script_unification_t *engine, uint32_t start, uint32_t end, decoder_script_range_t *gaps, size_t max_gaps, size_t *gap_count)

STRICT mode resolves overlaps by priority (PRIMARY < SECONDARY < LEGACY) with ucd_authoritative as a tie-breaker. LENIENT mode ignores priority and lets the first-added range cover any overlap. The pointer returned by decoder_script_unification_get_script() references engine-owned memory and is invalidated by subsequent add_range calls or by cleanup. The engine is not thread-safe.

Emoji (emoji.h)

Code Point Properties

  • bool decoder_is_emoji(uint32_t cp)
  • bool decoder_is_emoji_presentation(uint32_t cp)
  • bool decoder_is_emoji_modifier(uint32_t cp)
  • bool decoder_is_emoji_modifier_base(uint32_t cp)
  • bool decoder_is_emoji_component(uint32_t cp)
  • bool decoder_is_extended_pictographic(uint32_t cp)
  • bool decoder_is_regional_indicator(uint32_t cp)
  • bool decoder_is_emoji_variation_selector(uint32_t cp)
  • bool decoder_is_skin_tone_modifier(uint32_t cp)
  • decoder_emoji_skin_tone_t decoder_get_skin_tone(uint32_t cp)

Sequence Parsing

  • int decoder_next_emoji_sequence(const uint32_t *text, size_t len, size_t start_pos, decoder_emoji_sequence_info_t *info)
  • bool decoder_is_valid_flag_sequence(const uint32_t *text, size_t len)
  • bool decoder_is_keycap_sequence(const uint32_t *text, size_t len)
  • size_t decoder_count_emoji(const uint32_t *text, size_t len)

Security (security.h)

Identifier Validation

  • bool decoder_is_identifier_start(uint32_t cp)
  • bool decoder_is_identifier_continue(uint32_t cp)
  • bool decoder_is_valid_identifier(const uint32_t *str, size_t len)
  • bool decoder_is_pattern_syntax(uint32_t cp)
  • bool decoder_is_pattern_whitespace(uint32_t cp)
  • bool decoder_is_restricted_identifier_start(uint32_t cp)
  • bool decoder_is_restricted_identifier_continue(uint32_t cp)

Confusables and Skeletons

  • bool decoder_is_confusable(uint32_t a, uint32_t b)
  • decoder_confusable_type_t decoder_get_confusable_type(uint32_t a, uint32_t b)
  • int decoder_check_confusables(const uint32_t *s1, size_t len1, const uint32_t *s2, size_t len2)
  • int decoder_check_confusables_with_options(const uint32_t *s1, size_t len1, const uint32_t *s2, size_t len2, const decoder_security_options_t *options)
  • int decoder_get_skeleton(const uint32_t *src, size_t src_len, uint32_t *dst, size_t dst_capacity, size_t *dst_len)

Spoofing and Sanitization

  • bool decoder_is_suspicious(const uint32_t *str, size_t len)
  • bool decoder_is_spoofable(const uint32_t *str, size_t len)
  • bool decoder_is_highly_spoofable(const uint32_t *str, size_t len)
  • bool decoder_has_restricted_characters(const uint32_t *str, size_t len)
  • bool decoder_is_well_formed(const uint32_t *str, size_t len)
  • int decoder_sanitize(const uint32_t *src, size_t src_len, uint32_t *dst, size_t dst_capacity, size_t *dst_len, size_t *errors_removed)
  • int decoder_find_homoglyphs(uint32_t cp, decoder_homoglyph_result_t *result)
  • void decoder_homoglyph_result_free(decoder_homoglyph_result_t *result)

Character- and String-Level Checks

  • bool decoder_is_invisible(uint32_t cp)
  • bool decoder_is_restricted(uint32_t cp)
  • bool decoder_contains_invisible(const uint32_t *str, size_t len)
  • bool decoder_contains_mixed_numbers(const uint32_t *str, size_t len)
  • bool decoder_is_safe_filename(const uint32_t *str, size_t len)
  • bool decoder_is_safe_url(const uint32_t *str, size_t len)

High-Level UTF-8 Pipeline (decoder.h)

This layer is separate from the low-level modules. It is intended for UTF-8-first application code such as ingestion pipelines and CLI tools.

Lifecycle

  • void decoder_pipeline_init(void)
  • void decoder_pipeline_cleanup(void)
  • const char *decoder_pipeline_version(void)

UTF-8 Validation

  • decoder_utf8_validation_t
    • valid
    • codepoints
    • bytes_consumed
    • first_error_byte
    • error_byte_value
  • bool decoder_validate_utf8(const uint8_t *input, size_t input_len, decoder_utf8_validation_t *result)

UTF-8 Normalization and Sanitization

  • decoder_status_t decoder_normalize_text(const uint8_t *input, size_t input_len, decoder_normalization_form_t form, uint8_t *output, size_t output_cap, size_t *output_len)
  • bool decoder_text_is_normalized(const uint8_t *input, size_t input_len, decoder_normalization_form_t form)
  • decoder_sanitize_options_t
    • strip_control_chars
    • strip_format_chars
    • strip_replacement_chars
    • strip_private_use
    • strip_surrogates
    • normalize_whitespace
    • normalize_newlines
    • unwrap_lines
  • decoder_sanitize_options_t decoder_sanitize_defaults(void)
  • decoder_sanitize_options_t decoder_sanitize_preset_llm(void)
  • decoder_sanitize_options_t decoder_sanitize_preset_pdf(void)
  • decoder_status_t decoder_sanitize_text(const uint8_t *input, size_t input_len, const decoder_sanitize_options_t *options, uint8_t *output, size_t output_cap, size_t *output_len)

UTF-8 Case Folding

  • decoder_status_t decoder_casefold_text(const uint8_t *input, size_t input_len, uint8_t *output, size_t output_cap, size_t *output_len)

UTF-8 Grapheme Iteration

  • decoder_grapheme_t
    • byte_start
    • byte_end
    • codepoint_count
    • is_emoji
  • decoder_utf8_grapheme_iter_t
  • decoder_status_t decoder_utf8_grapheme_iter_init(decoder_utf8_grapheme_iter_t *it, const uint8_t *input, size_t input_len)
  • bool decoder_utf8_grapheme_iter_next(decoder_utf8_grapheme_iter_t *it, decoder_grapheme_t *grapheme)
  • void decoder_utf8_grapheme_iter_reset(decoder_utf8_grapheme_iter_t *it)
  • void decoder_utf8_grapheme_iter_free(decoder_utf8_grapheme_iter_t *it)
  • size_t decoder_count_graphemes_utf8(const uint8_t *input, size_t input_len)

Call decoder_utf8_grapheme_iter_free() if you initialized an iterator and do not exhaust it.

UTF-8 Security Scan

  • decoder_security_scan_t
    • has_mixed_scripts
    • has_confusable_chars
    • has_invisible_chars
    • has_bidi_override
    • has_restricted_chars
    • is_safe
    • threat_count
  • decoder_status_t decoder_security_scan(const uint8_t *input, size_t input_len, decoder_security_scan_t *result)
  • decoder_status_t decoder_skeleton_text(const uint8_t *input, size_t input_len, uint8_t *output, size_t output_cap, size_t *output_len)
  • bool decoder_are_confusable(const uint8_t *a, size_t a_len, const uint8_t *b, size_t b_len)

Text Analysis

  • decoder_text_stats_t
    • bytes
    • codepoints
    • graphemes
    • words
    • sentences
    • ascii_count
    • latin_count
    • emoji_count
    • combining_count
    • whitespace_count
    • control_count
    • punctuation_count
    • ascii_ratio
    • combining_ratio
  • decoder_status_t decoder_analyze_text(const uint8_t *input, size_t input_len, decoder_text_stats_t *stats)

Streaming

  • decoder_stream_config_t
    • form
    • case_fold
  • decoder_stream_config_t decoder_stream_defaults(void)
  • decoder_stream_t
  • decoder_status_t decoder_stream_open(decoder_stream_t *stream, const decoder_stream_config_t *config)
  • decoder_status_t decoder_stream_feed(decoder_stream_t *stream, const uint8_t *chunk, size_t chunk_len, uint8_t *output, size_t output_cap, size_t *output_len)
  • decoder_status_t decoder_stream_finish(decoder_stream_t *stream, uint8_t *output, size_t output_cap, size_t *output_len)

Full Pipeline Convenience API

  • decoder_pipeline_config_t
    • norm_form
    • case_fold
    • sanitize
    • security_scan
    • sanitize_opts
  • decoder_pipeline_config_t decoder_pipeline_defaults(void)
  • decoder_pipeline_config_t decoder_pipeline_preset_llm(void)
  • decoder_pipeline_result_t
    • status
    • security
    • output_len
    • was_normalized
  • decoder_status_t decoder_pipeline_process(const uint8_t *input, size_t input_len, const decoder_pipeline_config_t *config, uint8_t *output, size_t output_cap, decoder_pipeline_result_t *result)

Parallel Operations (parallel.h)

  • void decoder_parallel_set_threads(int num_threads)
  • int decoder_parallel_get_threads(void)
  • void decoder_parallel_to_upper(const uint32_t *input, size_t len, uint32_t *output)
  • void decoder_parallel_to_lower(const uint32_t *input, size_t len, uint32_t *output)
  • int decoder_parallel_normalize_nfc(const uint32_t *input, size_t len, uint32_t *output, size_t cap, size_t *output_len)
  • int decoder_parallel_normalize_nfd(const uint32_t *input, size_t len, uint32_t *output, size_t cap, size_t *output_len)
  • int decoder_parallel_get_skeleton(const uint32_t *input, size_t len, uint32_t *output, size_t cap, size_t *output_len)
  • void decoder_parallel_for(size_t count, void (*worker)(size_t start, size_t end, void *ctx), void *ctx)

These helpers are intended for native builds. The public header documents them as unavailable in single-threaded WASM builds.

Minimal Examples

Low-Level Grapheme Boundaries

#include <decoder.h>
#include <stdio.h>

int main(void) {
    decoder_init();

    const uint32_t text[] = { 'A', 0x0308, 'B' };
    size_t boundaries[8];
    size_t count = 0;

    if (decoder_find_grapheme_boundaries(text, 3, boundaries, 8, &count) == DECODER_SUCCESS) {
        printf("boundary count: %zu\n", count);
    }

    decoder_cleanup();
    return 0;
}

UTF-8 Pipeline

#include <decoder.h>

int main(void) {
    decoder_pipeline_init();

    const uint8_t input[] = "Straße";
    uint8_t output[64];
    size_t output_len = 0;

    decoder_pipeline_config_t cfg = decoder_pipeline_preset_llm();
    decoder_pipeline_result_t result;

    decoder_pipeline_process(input, sizeof(input) - 1, &cfg,
                             output, sizeof(output), &result);

    decoder_pipeline_cleanup();
    return 0;
}