CATEGORY
Text processing 4,995 crates
Data as of 2026-07-25 (crates.io database dump, timestamp 2026-07-25T02:00:36Z). Source: crates.io · db-dump.tar.gz · methodology & corrections.
english-tokenizerRegex-driven tokenizer that splits a sentence and type-tags each token (word, number, emoji, url, email, hashtag, mention, time,…40MIT
unicode-match-property-ecmascriptResolve a Unicode property name or alias to its canonical name for ECMAScript RegExp property escapes40MIT
csvprettyA command-line tool that formats CSV input into tables with Unicode box-drawing characters39MIT
strip-codeblocksA Rust library to strip markdown code blocks from text, preserving only the inner content39MIT
vagusLocal-first PARA second brain: hybrid full-text + semantic search over a Markdown vault39MIT
greek-utilsGreek transliteration, romanization, diacritics stripping and stopword removal39MIT
duallityLevenshtein automata as lling-llang WFSTs: composition adapters bridging liblevenshtein fuzzy matching and the lling-llang…39Apache-2.0
docxmlCreate and edit .docx files — a python-docx for Rust, built on a lossless XML tree for full round-trip fidelity39MIT OR Apache-2.0
ragrsFast local RAG in Rust. Index, query, verify.39MIT
redos-detectorstatic ReDoS vulnerability analysis of ECMA-262 regex patterns39MIT
spellchkA blazingly fast spellchecker CLI for any text file39MIT OR Apache-2.0
perl-ts-heredoc-analysisStandalone heredoc analysis tools for Perl parsing39MIT OR Apache-2.0
neco-syntax-textmateTextMate-style syntax loading and tokenization on top of syntect39MIT
newsfreshCLI and library for querying, filtering, and analyzing GDELT Global Knowledge Graph (GKG) v2.1 data — the world's largest open…39MIT
ansimakeQuickly convert pixel images of ANSI art created with AI to actual ANSI art39MIT
asimov-readwise-moduleASIMOV module.39Unlicense
utokenizerCLI tool for building a local model-tokenizer registry and counting input tokens across model families.39MIT
omni-mdxA highly secure, DoS-resistant MDX parser and OCP binary protocol engine.39MIT
onnx-genaiRust inference runtime for generative AI models on ONNX Runtime39MIT
arc-eagerArc-eager transition-based dependency parser using the MaltParser feature template39MIT
skill-extractorExtract skills from job postings and resumes — 32K-skill gazetteer + MiniLM embeddings + MLP context classifier (73% F1, trained…39MIT
brief-coreCompiler library for the Brief markup language: lexer, parser, AST, HTML/LLM emitters, formatter, and Markdown-to-Brief converter.39MIT
koreConvert JSON values into .kore format — a compact, human-readable data language. also for asura lang39MIT
graphql-strip-sensitive-literalsredact PII literals from a GraphQL query AST before logging/usage reporting39MIT
epub-stackEPUB Stack: a modern Rust implementation of the EPUB standard. (Name reserved, API under active design.)38MIT OR Apache-2.0
cvxtractLLM-powered structured extraction from CVs/resumes — PDF, DOCX, HTML, TXT input; typed Rust structs output.38MIT
danh-ngonKho danh ngôn song ngữ Việt–Anh (30.000+ câu song ngữ, 767.000+ kho đầy đủ) đóng gói sẵn, không cần mạng. Tìm theo chủ đề, tác…38MIT
formelockModern, high-performance document generation engine38Apache-2.0
html-entity-fixDecode HTML entities (& < > " ' numeric refs) that LLMs sometimes emit in plain-text output. Zero deps.38MIT OR Apache-2.0
neco-fuzzyMinimal fuzzy score core for commands, paths, and short identifiers38MIT
mq-openGraphical previewer for mq38MIT
asimov-anthropic-moduleASIMOV Anthropic module.38Unlicense
libphoneticsDatatypes for phones; sound change algorithms.38AGPL-3.0-only
zgSmall query-normalization helpers for the zg search tooling.38MIT OR Apache-2.0
slash-tokensToken budget estimation in 4.8 KB WASM. Sub-millisecond. Zero dependencies.38MIT
text-word-countcount words and characters in HTML/rich text38MIT
docrafter-fontTrueType font parsing, metrics, and PDF embedding for docrafter38MIT OR Apache-2.0
franken_markdownPure-Rust, dependency-lean, ultra-fast Markdown -> beautiful all-in-one HTML & tiny optimized PDF (library + single-binary CLI:…38non-standard
dfa-compilerCompile a regex-like DSL over integer symbol sequences into a DFA and scan for tagged matches38MIT
elephant-ladderElephant Ladder: a Rust-based reading system. (Name reserved.)37MIT OR Apache-2.0
pterPlain Text Email Renderer — convert HTML email bodies into readable markdown37MIT
llm-textA Rust library for processing text for LLM consumption37Apache-2.0
papercut-enginePapercut is an EPUB rendering and pagination engine for Rust. (Early pre-alpha, API under active design.)37MIT OR Apache-2.0
orion-coreBackend-agnostic agent harness for local LLM inference37MIT
uniko-apiPublic facade: builders, re-exports, no logic37Apache-2.0
anofox-ml-textText feature extraction (CountVectorizer, TfidfVectorizer, HashingVectorizer) for anofox-ml37MIT OR Apache-2.0
easy_strfmtA string formatting crate for maps and custom types37MIT OR Apache-2.0
promptchineA simple prompt refinement and caching library using LLMs37MIT
fontcoreLoad fonts, select faces, shape text, and export SVG in Rust.37MIT
cp-desktopCanon node ingestion pipeline — local-first AI knowledge management with anonymous sync37AGPL-3.0
molten_emberRender Markdown beautifully in the terminal 🔥37MIT OR Apache-2.0
rus-torchA comprehensive deep learning framework in Rust, merging core, nn, vision, text, and wasm37MIT OR Apache-2.0
file-ingestParse file bytes into a canonical document structure for AI and downstream processing.37MIT
asimov-gemini-moduleASIMOV Gemini module.37Unlicense
term-gptA fast, colorful ChatGPT CLI for your terminal!37MIT
geulbus-core날개셋(nalgaeset) 입력 설정을 해석하는 한글 조합 엔진 (ibus 비의존 순수 라이브러리)37MIT OR Apache-2.0
count-wordscount words in mixed CJK, Latin, and multilingual text37MIT
regex-rangecompile a numeric min/max range into the smallest matching regex source string37MIT
raster_fontA format for authoring and using image-backed fonts36MIT OR Apache-2.0
regex-specificityA heuristic-based crate to calculate the specificity of a regular expression pattern against a specific string.36MIT
panini-lang-langsLanguage-specific definitions and presets for the Panini linguistic feature extraction framework36MIT
nepali-phoneParse and validate Nepali phone numbers (mobile and landline).36MIT
paperdownA fast CLI tool to batch convert PDFs into Markdown using GLM-OCR.36MIT
brace-expansionBash-style brace expansion: a{b,c}d -> [abd, acd], {1..3} -> [1,2,3]. A faithful port of the brace-expansion npm package. Zero…36MIT OR Apache-2.0
sanitize-engineDeterministic one-way data sanitization engine36Apache-2.0
liepressA Markdown to PDF/SVG/PNG converter with CSS styling support36MIT OR Apache-2.0
textsanity-corePure-Rust core for textsanity: unicode/whitespace/encoding cleanup.36MIT OR Apache-2.0
popsam-cliCLI for AI-assisted selection of semantically representative texts36MIT
toklab-corePure-Rust core for toklab: bulk tokenizer + counter for OpenAI BPE encodings.36MIT OR Apache-2.0
kcode-web-fetchBounded public-web fetching with SSRF protection and readable text extraction36MIT
mdcliawesome-md terminal CLI/TUI (the `mdtool` binary): renders and live-edits markdown on macOS, Linux, and WSL.36MIT OR Apache-2.0
rd2qmd-sourceShared Rd source parsing facade for rd2qmd36MIT
expand-rangegenerate numeric/alphabetic ranges with steps and zero-padding for brace expansion36MIT
mdbook-obsidian-formatConvert Obsidian wiki links into mdBook-compatible Markdown36MIT OR Apache-2.0
scrybe-ratatuiRender a scrybe-core Markdown document as styled ratatui text — an embeddable Markdown view for any ratatui app36Apache-2.0
tokelang-coreCompression engine for Tokelang Lite: English-prompt compression middleware over standard tokenizers, with a content-recall…35Apache-2.0
docx-review-cliCLI for extracting review-oriented structured data from DOCX files35MIT
tokitBlazing fast parser combinators: parse-while-lexing (zero-copy), deterministic LALR-style parsing, no backtracking. Flexible…35MIT OR Apache-2.0
triplets-srd-sourceExperimental simd-r-drive integration for the triplets data pipeline framework.35MIT OR Apache-2.0
discountMarkdown parser and display library for terminals35Apache-2.0
neco-editorUmbrella crate for editor runtime primitives with a unified text buffer35MIT
jnanaJnana — the foundation of knowing. Unified knowledge system for AGNOS35GPL-3.0-only
sdaStructured Data Algebra command-line interface for evaluating, checking, and formatting SDA programs over JSON input.35GPL-3.0-or-later
invlex-cliCLI tool for inverse lexicographic (a tergo) sorting; installs the `invlex` binary35MIT OR Apache-2.0
markalignCompare Markdown documents at the syntax level35BSD-3-Clause
js_ergoErgonomic, JavaScript-style string helpers for Rust (padStart and friends).35MIT
bridgexOpen-source desktop app for converting files to Markdown, built in Rust with Freya and Markitdown.35GPL-3.0-only
secretsniff-corePure-Rust core for secretsniff: source-code secret scanner.35MIT OR Apache-2.0
stream-rsZero-dependency, spec-compliant streaming toolkit for LLM responses (SSE, incremental JSON, OpenAI/Anthropic delta accumulators).35MIT
snipsplit-corePure-Rust core for snipsplit: token-aware text chunker for RAG ingestion.35MIT OR Apache-2.0
papyrust-cliBuild publication-quality EPUB and print-ready PDF from a folder of Markdown.35MIT
diffwtf-coreFast Myers diff engine with intra-line refinement — powers diff.wtf35MIT
ipsortVersitile ip address sorting tool35GPL-3.0-or-later
winditGeneric windowed-sequence processing: chunk, pad/mask, aggregate, smooth, segment — for embeddings, VAD, and ASR.35MIT OR Apache-2.0
case_conv_macrosMacros to convert identifiers and string literals to strings in a different case style.34MIT
agentrootFast local semantic search for codebases and knowledge bases with AI-powered features34MIT
memora-cliMemora: catch your AI citing sources that don't say what it claims. Local verification layer for AI memory.34Apache-2.0
tok3niz3r-cliCommand-line tool to train, run, and inspect byte-level BPE tokenizers. Installs the `tok3` binary.34MIT
mdbook-tsAn mdBook preprocessor that uses tree-sitter to extract code snippets from source files34MIT
nerdleA macro-powered compile-time nerd-font code point resolver.34MIT
ytx-cliExtract YouTube transcripts from the terminal. Pipe-friendly, no API key needed.34MIT
defuddle-rsA Rust library for extracting main content and metadata from HTML web pages34MIT OR Apache-2.0
missionRobust HTML parser and CSS selector engine. Won't crash on bad HTML.34MIT OR Apache-2.0
rune-chain-splitterText splitters for LLM context windows: character, token, markdown, and recursive strategies34MIT
maskprompt-corePure-Rust core for maskprompt: PII redaction for LLM prompts.34MIT OR Apache-2.0
ogham-serverEmbeddable HTTP server for the Ogham context engineering SDK34Apache-2.0
rostdownA kramdown-compatible Markdown renderer (GFM-flavored subset) producing byte-identical HTML; Rost is German for rust, the way…34MIT OR Apache-2.0
glyphrush-lopdfDependency-light lopdf extraction backend for the Glyphrush PDF parser34MIT
picomdA tiny, offline, GitHub-fidelity Markdown previewer (no editor).33MIT
capsaA compact, lightweight library for embedding-based document storage and retrieval33MIT
tamil-yaappu-analyzerTamil prosody analyzer and classifier for verse compositions33Apache-2.0
easi-publishCompile Typst templates with data into PDF or HTML documents.33Apache-2.0
window-enumerator-formatterA powerful formatting library for window information with multiple output formats (JSON, YAML, CSV, Table) and template support33MIT OR Apache-2.0
karmisttodoist, but for gigachads33GPL-3.0-or-later
mos-lspLanguage server for Mosaic (manifest §17).33MIT
kizameKizaMe (刻め!) - CLI for MeCrab morphological analyzer and data pipeline33MIT OR Apache-2.0
mecab-ko-dict-syncKorean National Institute dictionary API client for MeCab-Ko33MIT OR Apache-2.0
rust365A fast, dependency-free Microsoft Word .docx to HTML converter, in Rust.33MIT
webfetch-coreShared primitives: compression, token budgeting, and reference-style URLs33MIT
salary-parser-jpParse Japanese job-posting salary text (月給21万円〜26万円) into typed yen amounts: hourly, daily, monthly, yearly. Normalizes…33Apache-2.0
mem1-coreCore types for mem1, a long-term memory system for AI conversations33MIT
khmer-tokenizer-coreFast, dependency-free Khmer word segmentation (syllable-aware longest match).33MIT OR Apache-2.0
nacre-coreLLM extraction pipeline for agent memory — Graphiti's pipeline ported to Rust on grit32Apache-2.0
ikigai-textUnix-like text endpoints (urn:text:wc/head/tail/grep/sort/uniq/nl/rev) for ikigai — pure, cacheable, wasm-clean pipeline citizens…32MIT OR Apache-2.0
lexaLexa CLI: hybrid local search (BM25 + binary-quantized Matryoshka KNN + cross-encoder rerank) over arbitrary file trees. `lexa…32MIT OR Apache-2.0
tuika-codeformattersTree-sitter syntax highlighting for tuika's CodeBlock and Markdown components — a ready-made Highlighter implementation.31MIT
tokelang-mecModel-Emulation Compression — heuristic, faithfulness-gated prompt compressor (clean-room).31Apache-2.0
mdyaLocal Markdown retrieval primitive (BM25 + on-device vector + MCP).31MIT OR Apache-2.0
agentic-evalEvaluate programs, CLI commands, programming languages, AI frameworks, and VM/sandbox systems for agentic AI use across four axes…31AGPL-3.0-or-later
lectito-mcpLocal stdio MCP server for Lectito article search and reading.31MPL-2.0
molten_sigilHuman-readable ANSI escape sequences for terminal styling ✨31MIT OR Apache-2.0
mosCommand-line interface for the Mosaic typesetting engine (manifest §15.1).31MIT
markmaidFramework-agnostic Markdown rendering engine in pure Rust: hand-written GFM-subset parser, document layout to plain positioned…31GPL-3.0-or-later
picomatch-rsRust glob matching core for the picomatch-rs workspace.31MIT
roff-cliSkillful man page to JSON/Markdown converter - human readable, AI-friendly31MIT
fill-rangeFill in a range of numbers or letters: fill("1","5") -> [1,2,3,4,5], fill("a","e") -> [a,b,c,d,e]. A faithful port of the…31MIT OR Apache-2.0
slice-ansiSlice a string by terminal display column, preserving ANSI styling and wide (CJK) characters. A Rust take on Node's slice-ansi.31MIT OR Apache-2.0
lexa-mcprmcp stdio MCP server for the Lexa hybrid retrieval engine. Exposes `search_files`, `index_path`, `list_indexed_paths`,…31MIT OR Apache-2.0
hippmem-modelModel backends for HIPPMEM — Embedder, Extractor, Reranker, Summarizer traits and registry31Apache-2.0
wmedrano-test-crateSubsetting a font file according to provided input.31MIT OR Apache-2.0
sdsconv-coreConvert chemical safety SDS documents (PDF/DOCX) ↔ MHLW/JIS Z 7253 standard JSON via LLM. Supports Claude, GPT, Gemini.…31MIT OR Apache-2.0
forbidden-regexRestricted byte-oriented regex engine for line-at-a-time secret scanning, with counting, product, and DFA back-ends, set-level…31LGPL-3.0-or-later
flux-md-coreIncremental, streaming-aware markdown parser with speculative closure31MIT
cli-truncateTruncate a string to a terminal display width with an ellipsis — wide-character (CJK) and ANSI-escape aware.31MIT OR Apache-2.0
searchexA searchable-model layer for Rust: make a type searchable, keep the index in sync, and search with real relevance ranking — over…30MIT
brevisAn XML processor for a more comfortable light markup syntax30Apache-2.0
tulisp-fmtSource code formatter for tulisp / Emacs Lisp.30GPL-3.0-only
hy-mtA lightweight machine translation inference library for Tencent Hunyuan MT models30MIT
md2pdf-rsA CLI tool to convert Markdown to PDF using Typst30MIT
fast-unescape'unescapes' a escaped string with escape sequences into literal one30LGPL-3.0-or-later
mdsqlSQL queries for markdown tables30GPL-2.0
detect-newlineDetect and normalize line endings (LF, CRLF, CR) — a zero-dependency, no_std toolkit. A Rust take on Node's detect-newline.30MIT OR Apache-2.0
indentkitDetect a string's indentation (tabs vs spaces + size) and re-indent / convert between styles. Like Node's detect-indent.…30MIT OR Apache-2.0
tf-idf-matcherApproximate string matching using n-gram TF-IDF vectorization and cosine similarity30MIT
lexrs-serverProduction HTTP server for the lexrs lexicon library30MIT
mem1-dbDatabase layer and memory service for mem130MIT
siarA grep-like, pure Rust PDF text searcher30AGPL-3.0-only
win-text-injectDeliver text into the focused Windows application correctly: clipboard privacy formats, modifier sanitization, integrity-level…30MIT OR Apache-2.0
hangeul_jamo_rsA high-performance Korean Hangul syllable and jamo manipulation library. included Python bindings.29MIT
flash-fuzzy-coreHigh-performance fuzzy search using Bitap algorithm with bloom filter pre-filtering. Zero dependencies, no_std compatible.29MIT
memory-wikiA local-first, semantic knowledge base and MCP server for LLMs.29MIT
agentroot-cliFast local semantic search for codebases and knowledge bases with AI-powered features29MIT
prettycharsUnicode text styling and named glyph lookup with zero runtime overhead29MIT OR Apache-2.0
rnk-style-coreShared style primitives for rnk crates29MIT
merge-engineA non-LLM merge conflict resolver using structured merge, Version Space Algebra, and search-based techniques29MIT
ai-translator基于 AI 的多语言文本翻译工具,支持自定义提示词29Apache-2.0
man_parserA Rust roff parser for converting man pages to JSON/Markdown29MIT
smartypantsTranslate plain ASCII punctuation into smart typographic punctuation — curly quotes, em/en-dashes and ellipses. A…29MIT OR Apache-2.0
to-regex-rangeGenerate a compact regex that matches an integer range: (1, 99) -> [1-9]|[1-9][0-9]. A faithful port of the to-regex-range npm…29MIT OR Apache-2.0
nibdexnibdex — derived MCP index over a workspace's source code, git commits, design docs, memory, and AI session history29MIT
docx-to-mdParse Microsoft Word files (.docx) and OpenDocument Text file (.odt) into Markdown (.md)29MIT OR Apache-2.0
badwords-coreCore profanity filter logic - normalization, transliteration, homoglyphs28MIT
ghost-libGhost Librarian — ultra-lightweight local-LLM RAG engine with Context Distillation28MIT
vn-nlpVietnamese NLP library — tokenization, normalization, segmentation28MIT OR Apache-2.0
ai-marketing-campaign-optimizerAI Marketing Campaign Optimizer - Multi-language toolkit for optimizing AI-powered marketing campaigns with content analysis,…28MIT
cronifyA tool to convert natural language time expressions into cron syntax.28MIT
org-toolsUnified CLI for org-mode: lint, format, query, clock, export28GPL-3.0-or-later
docx_basicA minimal, style-first DOCX generation library28MIT
zwusZero Width Unicode Steganography — hide text in invisible characters.28WTFPL
papyrus-cliCommand-line tool for PDF-to-Markdown conversion with smart heading detection, bold/italic extraction, and CommonMark output.…28MIT OR Apache-2.0
chordsketch-convertChordPro ↔ iReal Pro format-conversion bridge (trait scaffold)28MIT
rjtd-coreLow-level parsers and diagnostics for Ichitaro JTD compound documents.28Apache-2.0
mf2_i18n_coreCore types and runtime primitives for Unicode MessageFormat v2 (MF2).28MIT OR Apache-2.0
kohagiLocal sentence embeddings for Ruri v3 and other ModernBERT models: JSONL in, vectors out. Pure Rust, bounded memory.28MIT
femindPluggable, feature-gated memory engine for AI agent applications27MIT OR Apache-2.0
tpt-jinja-chatPure-Rust Jinja2 subset parser for LLM chat templates — zero dependencies.27MIT OR Apache-2.0
md-rename-coreCore library for renaming Markdown files and rewriting links.27MIT
context-compressorDynamic token budget context reducer with regex sanitizers.27Apache-2.0
drapeANSI-aware text wrapping built on top of textwrap27MIT/Apache-2.0
pleias-stratum-clientOfficial Rust client for the Stratum document-chunking API.27Apache-2.0
mdbook-listingsManaged code listings for mdbook: inline callouts, freezing, and verification27MIT
banana-i18nA Rust library for internationalization (i18n) with MediaWiki-style message formatting and localization.27MIT
triplets-offline-embedderOffline teacher-embedding precompute pipeline for the triplets data framework.27MIT OR Apache-2.0
elizaos-plugin-eliza-classicClassic ELIZA pattern matching plugin for elizaOS - no LLM required27MIT
typf-osPlatform-native linra text rendering dispatcher27Apache-2.0
tokidToken-native IDs for LLM-facing systems27ISC
verso-readerA terminal EPUB reader with vim navigation, a Kindle-style library, and Markdown highlight export27MIT OR Apache-2.0
loggrepA smarter log parser with color-coded severity, time filtering, regex matching, and stats27MIT
pdfoxA pure-Rust PDF library — create, parse, and render PDF documents with zero C dependencies27MIT
langrFast, low-memory single-purpose language detector using an LLM subword vocab and token->language posteriors.27MIT