CATEGORY
Text processing 4,995 crates
Data as of 2026-07-25 (crates.io database dump, timestamp 2026-07-25T02:00:36Z). Source: crates.io · db-dump.tar.gz · methodology & corrections.
lingua-estonian-language-modelThe Estonian language model for Lingua, an accurate natural language detection library1,395,349Apache-2.0
lingua-bengali-language-modelThe Bengali language model for Lingua, an accurate natural language detection library1,393,736Apache-2.0
lingua-serbian-language-modelThe Serbian language model for Lingua, an accurate natural language detection library1,393,700Apache-2.0
lingua-tagalog-language-modelThe Tagalog language model for Lingua, an accurate natural language detection library1,393,344Apache-2.0
lingua-albanian-language-modelThe Albanian language model for Lingua, an accurate natural language detection library1,393,082Apache-2.0
lingua-bosnian-language-modelThe Bosnian language model for Lingua, an accurate natural language detection library1,392,432Apache-2.0
lingua-georgian-language-modelThe Georgian language model for Lingua, an accurate natural language detection library1,392,092Apache-2.0
lingua-armenian-language-modelThe Armenian language model for Lingua, an accurate natural language detection library1,391,501Apache-2.0
lingua-kazakh-language-modelThe Kazakh language model for Lingua, an accurate natural language detection library1,391,214Apache-2.0
lingua-tamil-language-modelThe Tamil language model for Lingua, an accurate natural language detection library1,390,550Apache-2.0
lingua-macedonian-language-modelThe Macedonian language model for Lingua, an accurate natural language detection library1,390,305Apache-2.0
lingua-gujarati-language-modelThe Gujarati language model for Lingua, an accurate natural language detection library1,390,010Apache-2.0
lingua-punjabi-language-modelThe Punjabi language model for Lingua, an accurate natural language detection library1,389,540Apache-2.0
lingua-telugu-language-modelThe Telugu language model for Lingua, an accurate natural language detection library1,389,348Apache-2.0
lingua-somali-language-modelThe Somali language model for Lingua, an accurate natural language detection library1,389,116Apache-2.0
lingua-swahili-language-modelThe Swahili language model for Lingua, an accurate natural language detection library1,388,920Apache-2.0
lingua-marathi-language-modelThe Marathi language model for Lingua, an accurate natural language detection library1,388,867Apache-2.0
lingua-belarusian-language-modelThe Belarusian language model for Lingua, an accurate natural language detection library1,356,369Apache-2.0
lingua-azerbaijani-language-modelThe Azerbaijani language model for Lingua, an accurate natural language detection library1,356,170Apache-2.0
lingua-irish-language-modelThe Irish language model for Lingua, an accurate natural language detection library1,355,854Apache-2.0
lingua-afrikaans-language-modelThe Afrikaans language model for Lingua, an accurate natural language detection library1,355,177Apache-2.0
lingua-esperanto-language-modelThe Esperanto language model for Lingua, an accurate natural language detection library1,355,010Apache-2.0
lingua-basque-language-modelThe Basque language model for Lingua, an accurate natural language detection library1,354,708Apache-2.0
lingua-welsh-language-modelThe Welsh language model for Lingua, an accurate natural language detection library1,354,438Apache-2.0
lingua-zulu-language-modelThe Zulu language model for Lingua, an accurate natural language detection library1,354,144Apache-2.0
lingua-xhosa-language-modelThe Xhosa language model for Lingua, an accurate natural language detection library1,354,002Apache-2.0
lingua-shona-language-modelThe Shona language model for Lingua, an accurate natural language detection library1,353,963Apache-2.0
lingua-maori-language-modelThe Māori language model for Lingua, an accurate natural language detection library1,353,816Apache-2.0
lingua-ganda-language-modelThe Ganda language model for Lingua, an accurate natural language detection library1,353,738Apache-2.0
lingua-tswana-language-modelThe Tswana language model for Lingua, an accurate natural language detection library1,353,413Apache-2.0
lingua-sotho-language-modelThe Sotho language model for Lingua, an accurate natural language detection library1,353,392Apache-2.0
lingua-yoruba-language-modelThe Yoruba language model for Lingua, an accurate natural language detection library1,352,987Apache-2.0
lingua-tsonga-language-modelThe Tsonga language model for Lingua, an accurate natural language detection library1,352,679Apache-2.0
roman-numerals-rsManipulate well-formed Roman numerals1,350,1760BSD OR CC0-1.0
lingua-latin-language-modelThe Latin language model for Lingua, an accurate natural language detection library1,348,619Apache-2.0
imperativeCheck for imperative mood in text1,340,725MIT OR Apache-2.0
lindera-ko-dicA Korean morphological dictionary for ko-dic.1,307,405MIT
editOpen a file in the default text editor1,282,844CC0-1.0
flexstrA flexible, simple to use, clone-efficient string type for Rust1,253,278MIT OR Apache-2.0
new_string_templateSimple Customizable String-Templating Library for Rust.1,220,384MIT
rapidfuzzrapid fuzzy string matching library1,208,392MIT
detoneDecompose Vietnamese tone marks1,149,960Apache-2.0 OR MIT
wcharProcedural macros for compile time UTF-16 and UTF-32 wide strings.1,149,668MIT OR Apache-2.0
lindera-cc-cedictA Chinese morphological dictionary for CC-CEDICT.1,133,430MIT
convert_case_extrasExtra features for convert_case1,125,610MIT
xlsxwriterWrite xlsx file with number, formula, string, formatting, autofilter, merged cells, data validation and more.1,118,937Apache-2.0
charabiaA simple library to detect the language, tokenize the text and normalize the tokens1,082,410MIT
lindera-unidicA Japanese morphological dictionary for UniDic.1,075,257MIT
llm-tokenizerLLM tokenizer library with caching and chat template support1,057,141Apache-2.0
granit-parserA YAML parser with comment and style support, written in pure Rust1,043,939MIT OR Apache-2.0
reasoning-parserParser for AI model reasoning/thinking outputs (chain-of-thought, etc.)1,009,775Apache-2.0
synopticA simple, low-level, syntax highlighting library with unicode support996,322MIT
sanitizerA collection of methods and macros to sanitize struct fields.963,076MIT
slugifyMacro for flexible slug generation958,386MIT
yaml_parserSemi-tolerant YAML concrete syntax tree parser.956,646MIT
lindera-ipadic-neologdA Japanese morphological dictionary for IPADIC NEologd.926,995MIT
segtokSentence segmentation and word tokenization tools926,893MIT
typos-cliSource Code Spelling Correction916,651MIT OR Apache-2.0
nnsplitA tool to split text using a neural network. For sentence boundary detection, compound splitting and more.884,592MIT
umya-spreadsheetumya-spreadsheet is a library written in pure Rust to read and write xlsx file.872,163MIT
html-to-markdown-rsHigh-performance HTML to Markdown converter using the astral-tl parser. Part of the Xberg ecosystem.842,061MIT
cedictParser for the CC-CEDICT Chinese-English Dictionary830,831MIT
lindera-coreA morphological analysis library.830,396MIT
sentencexSentence segmentation library with wide language support optimized for speed and utility.828,332MIT
daggrsA fast Double-Array Aho-Corasick implementation for multi-pattern matching821,266MIT OR Apache-2.0
lindera-ipadic-builderA Japanese morphological dictionary builder for IPADIC.808,585MIT
unicode-canonical-combining-classFast lookup of the Canonical Combining Class property794,983Apache-2.0
glyph-namesMapping of characters to glyph names according to the Adobe Glyph List Specification782,856BSD-3-Clause
lindera-unidic-builderA Japanese morphological dictionary builder for UniDic.777,627MIT
simple-loggingA simple logger for the log facade773,725BSD-3-Clause
pagerHelps pipe your output through an external pager772,562Apache-2.0/MIT
lindera-decompressA morphological analysis library.772,223MIT
lindera-ko-dic-builderA Korean morphological dictionary builder for ko-dic.770,592MIT
lindera-cc-cedict-builderA Chinese morphological dictionary builder for CC-CEDICT.765,202MIT
precis-profilesImplementation of the PRECIS Framework: Preparation, Enforcement,
and Comparison of Internationalized Strings…751,596MIT/Apache-2.0
precis-corePRECIS Framework: Preparation, Enforcement, and Comparison of
Internationalized Strings in Application Protocols as defined
in…745,248MIT/Apache-2.0
precis-toolsTools and parsers to generate PRECIS tables from the Unicode Character Database (UCD)742,892MIT/Apache-2.0
ferris-saysA Rust flavored replacement for the classic cowsay734,643MIT OR Apache-2.0
harfbuzz-traitsRust Traits for the HarfBuzz text shaping engine720,907MIT OR Apache-2.0
allsortsFont parser, shaping engine, and subsetter for OpenType, WOFF, and WOFF2714,425Apache-2.0
sql-builderSimple SQL code generator.710,941MIT
decancerA library that removes common unicode confusables/homoglyphs from strings.710,545MIT
bk-treeA Rust BK-tree implementation671,494MIT
naturalPure rust library for natural language processing.649,075MIT
noto-sans-mono-bitmapProvides pre-rasterized characters from the "Noto Sans Mono" font in different sizes and font
weights for multiple unicode…644,700MIT
quoted-string-parserQuoted string parser for grammar defined in RFC3261641,176MIT/Apache-2.0
typos-dictSource Code Spelling Correction638,551MIT OR Apache-2.0
extractousExtractous provides a fast and efficient way to extract content from all kind of file formats including PDF, Word, Excel
CSV,…617,702Apache-2.0
textdistanceLots of algorithms to compare how similar two sequences are603,696MIT
unic-ucd-normalUNIC — Unicode Character Database — Normalization Properties567,912MIT/Apache-2.0
unic-ucd-hangulUNIC — Unicode Character Database — Hangul Syllable Composition & Decomposition562,494MIT/Apache-2.0
typosSource Code Spelling Correction554,046MIT OR Apache-2.0
gh-emojiConvert `:emoji:` to Unicode using GitHub's emoji names545,546MIT
lowchartsTool to draw low-resolution graphs in terminal537,355MIT
unic-ucd-ageUNIC — Unicode Character Database — Age531,703MIT/Apache-2.0
wana_kanaUtility library for checking and converting between Japanese characters - Kanji, Hiragana, Katakana - and Romaji526,241MIT
varcon-coreVarcon-relevant data structures522,603MIT OR Apache-2.0
typos-varsSource Code Spelling Correction521,222MIT OR Apache-2.0
lindera-ipadic-neologd-builderA Japanese morphological dictionary builder for IPADIC NEologd.515,962MIT
genpdfUser-friendly PDF generator written in pure Rust510,721Apache-2.0 OR MIT
sdAn intuitive find & replace CLI505,858MIT
unic-normalUNIC — Unicode Normalization Forms504,809MIT/Apache-2.0
dictgenCompile-time case-insensitive map492,451MIT OR Apache-2.0
etchNot just a text formatter, don't mark it down, etch it.480,593MIT
lindera-compressA morphological analysis library.479,085MIT
tiny-gradientMake your string colored in gradient477,776MIT
line-spanFind line ranges and jump between next and previous lines438,573MIT
unic-ucd-commonUNIC — Unicode Character Database — Common Properties437,979MIT/Apache-2.0
lindera-tokenizerA morphological analysis library.429,376MIT
simsearchA small in-memory fuzzy search index for embedded autocomplete and search suggestions.419,109MIT
pdf_oxideThe fastest Rust PDF library — 0.8ms mean, 5× faster than the industry leaders, 100% pass rate on 3,830 real-world PDFs. Text…399,518MIT OR Apache-2.0
harfbuzzHigh-level Rust bindings to the HarfBuzz text shaping engine390,948MIT OR Apache-2.0
markdown-genCrate for generating Markdown files380,774MIT
mdxjsCompile MDX to JavaScript in Rust.368,800MIT
qp-trieAn idiomatic and fast QP-trie implementation in pure Rust, written with an emphasis on safety.355,606MPL-2.0
json_to_tableA library for pretty print JSON as a table350,200MIT
bytelinesRead input lines as byte slices for high efficiency349,513MIT
epub-parserA Rust library for extracting metadata, table of contents, text, cover, and images from EPUB files.348,981MIT
omniparseA Rust toolkit for detecting and extracting metadata, text, and content from various file formats348,037MIT OR Apache-2.0
compact_bytesA memory efficient bytes container that transparently stores bytes on the stack, when possible333,037MIT/Apache-2.0
fontconfigSafe, higher-level wrapper around the Fontconfig library331,525MIT
whichlangA blazingly fast and lightweight language detection library for Rust.324,980MIT
ansi-escape-sequencesHigh-performance Rust library for detecting, matching, and processing ANSI escape sequences in terminal text with zero-allocation…322,205MIT OR Apache-2.0
wild-docYou can read and write data using XML and output various structured documents.You can also program using…320,305MIT/Apache-2.0
east-asian-widthDetermine the display width of Unicode characters in East Asian contexts317,192MIT OR Apache-2.0
string-widthAccurate Unicode string width calculation for terminal applications, handling emoji, East Asian characters, combining marks, and…316,167MIT OR Apache-2.0
wrap-ansiA high-performance, Unicode-aware Rust library for intelligently wrapping text while preserving ANSI escape sequences, colors,…315,256MIT OR Apache-2.0
linkcheckA library for extracting and validating links.308,473MIT OR Apache-2.0
ratatui-textarearatatui-textarea is a simple yet powerful text editor widget for ratatui.
Multi-line text editor can be easily put as part of…298,406MIT
toml-test-dataTOML test cases293,282MIT OR Apache-2.0
async-logAsync tracing capabilities for the log crate.292,682MIT OR Apache-2.0
xmldeclExtracts an encoding from an ASCII-based bogo-XML declaration in text/html in a Web-compatible way287,944Apache-2.0 OR MIT
toml-testVerify Rust TOML parsers285,555MIT OR Apache-2.0
bwrapA fast, lightweight, embedded systems-friendly library for
wrapping text.284,360MIT OR GPL-3.0-or-later
office_oxideThe fastest Office document processing library — DOCX, XLSX, PPTX, DOC, XLS, PPT279,963MIT OR Apache-2.0
jsonxfA fast JSON pretty-printer and minimizer.277,396MIT
toml-test-harnessCargo test harness for verifying TOML parsers275,780MIT OR Apache-2.0
harfbuzz_rsA high-level interface to HarfBuzz, exposing its most important functionality in a safe manner using Rust.269,333MIT
async-log-attributesProc Macro attributes for the async-log crate.267,933MIT OR Apache-2.0
vaporettoVaporetto: a pointwise prediction based tokenizer262,602MIT OR Apache-2.0
string_templateVery simple string template for Rust.260,960MIT/Apache-2.0
censorA simple text profanity filter248,787MIT
markdown-itRust port of popular markdown-it.js library.245,601MIT
slugify-rsA rust library to generate slugs from strings241,393non-standard
rewordProvides some utility functions for human-readable formatting of words.239,241MIT OR Apache-2.0
cropA pretty fast text rope237,983MIT
partial-json-fixerPartial JSON fixer fixes partial JSON236,182MIT
dom_smoothieA Rust crate for extracting relevant content from web pages234,620MIT
regex-cursorregex fork that can search discontiguous haystacks228,412MIT OR Apache-2.0
markdown-pppFeature-rich Markdown Parsing and Pretty-Printing library226,187MIT
merge-junitCLI utility to merge JUnit compliant XML documents into a single XML document.222,060Unlicense
tp-noteThis crate has moved to `tpnote`215,761MIT/Apache-2.0
mandownMarkdown to groff (man page) converter214,957Apache-2.0 OR MIT
stringsliceA collection of methods to slice strings based on character indices rather than bytes214,904MIT OR Apache-2.0
dedentProcedural macro for stripping indentation from multi-line string literals212,303MIT
unicode-ellipsisA crate to truncate Unicode strings to a certain width, automatically adding an ellipsis if the string is too long.212,058MIT OR Apache-2.0
str-utilsThis crate provides some traits to extend `[u8]`, `str` and `Cow<str>`.209,270MIT
lindera-tantivyLindera Tokenizer for Tantivy.208,044MIT
lindera-filterCharacter and token filters for Lindera.207,896MIT
ressA scanner/tokenizer for JS files207,121MIT
mdbook-svgbobSvgBob mdbook preprocessor which swaps code-blocks with neat SVG.202,725MPL-2.0
character_converterTurn Traditional Chinese script ot Simplified Chinese script and vice-versa and tokenize.202,480MIT
regex-macroA macro to generate a lazy regex expression200,388MIT OR Apache-2.0
rphoneticRust port of phonetic Apache commons-codec algorithms197,542Apache-2.0
xhtmlchardetCharacter set detection for XML and HTML197,214MIT
unicode-display-widthUnicode 15.1.0 compliant utility for determining the number of columns required to display an arbitrary string196,644MIT
simple_excel_writerSimple Excel Writer195,353Apache-2.0
tinyvec_stringtinyvec based string types193,482MIT OR Apache-2.0
mdcatcat for markdown: Show markdown documents in terminals192,707MPL-2.0 AND Apache-2.0
ruby_inflectorAdds String based inflections for Rust. Snake, kebab, camel, sentence, class, title and table cases as well as ordinalize,…191,488BSD-2-Clause
vectorscan-rs-sysNative bindings to the Vectorscan high-performance regex library186,588Apache-2.0 OR MIT
tectonicA modernized, complete, embeddable TeX/LaTeX engine. Tectonic is forked from the XeTeX
extension to the classic "Web2C"…186,369MIT
textwrap-macros-implSimple procedural macros to use textwrap utilities at compile time.183,698MIT
textwrap-macrosSimple procedural macros to use textwrap utilities at compile time.182,823MIT
lindera-analyzerA morphological analysis library.182,722MIT
web-rwkvAn implementation of the RWKV language model in pure WebGPU.179,781MIT OR Apache-2.0
sanitise-file-nameAn unusually flexible and efficient file name sanitiser177,000BlueOak-1.0.0 OR MIT OR Apache-2.0
vectorscan-rsErgonomic bindings to the Vectorscan high-performance regex library174,161Apache-2.0 OR MIT
symspellSpelling correction & Fuzzy search164,784MIT
unic-ucd-nameUNIC — Unicode Character Database — Name161,506MIT/Apache-2.0
orgA library for handling org-mode files.161,240MIT
pinotFast, high-fidelity OpenType parser.160,910MIT OR Apache-2.0
capitalizeChange first character to upper case and the rest to lower case, and other common alternatives158,620Unlicense
in_definiteGet the indefinite article ('a' or 'an') to match the given word. For example: an umbrella, a user.155,146MIT
unic-ucdUNIC — Unicode Character Database154,736MIT/Apache-2.0
uwuifyfastest text uwuifier in the west154,146MIT
diff-match-patch-rsThe fastest implementation of Myer's diff algorithm to perform the operations required for synchronizing plain text.153,517MIT OR Apache-2.0
fuzzydateA flexible natural language date parsing library153,079MIT
nlpruleA fast, low-resource Natural Language Processing and Error Correction library.152,226MIT OR Apache-2.0
lindera-dictionary-builderShared code for building Lindera dictionary files151,967MIT
neo_frizbeeFast typo-resistant fuzzy matching via SIMD smith waterman, similar algorithm to FZF/FZY151,931MIT
unic-ucd-caseUNIC — Unicode Character Database — Case Properties151,398MIT/Apache-2.0
codegenrsMoving code-gen our of build.rs150,007MIT OR Apache-2.0
aneubeck-daachorseDaachorse: Double-Array Aho-Corasick146,822MIT OR Apache-2.0
durstrA simple library for parsing human-readable strings into durations.144,908MIT
secularNo Diacr!143,350MIT
include-linesMacros for reading in the lines of a file at compile time143,050MIT OR Apache-2.0
fast2sA fast Traditional Chinese to Simplified Chinese conversion library. Built with FST, faster than most of other libraries.141,883MIT
unic-ucd-blockUNIC — Unicode Character Database — Unicode Blocks141,259MIT/Apache-2.0
tekken-rsRust implementation of Mistral Tekken tokenizer with audio support140,558Apache-2.0