The Principal Dev – Masterclass for Tech Leads

The Principal Dev – Masterclass for Tech Leads28-29 May

Join

anydoc

Crates.io npm PyPI License: MIT skills.sh

Fast Rust library that converts documents (Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF) into clean GitHub-Flavored Markdown. Includes bindings for Node.js and Python.

Built by Firecrawl to turn any office document into LLM-ready Markdown in single-digit milliseconds, with one consistent output no matter which format goes in. It powers Firecrawl Parse, so if you'd rather not run it yourself, the hosted API gives you the same conversion plus our OCR models for the scanned pages anydoc can't read on its own.

Quick start

Agent skill

anydoc ships as an Agent Skill, so your agent can read any document it runs into:

npx skills add firecrawl/anydoc

The skill teaches the agent to convert documents with the anydoc CLI. Works with Claude Code, Codex, Cursor, OpenCode, and any other compatible agent.

CLI

npx @firecrawl/anydoc report.docx               # Markdown to stdout
npx @firecrawl/anydoc slides.pptx -o slides.md  # or to a file
npx @firecrawl/anydoc - --format csv < data.csv # read stdin

npx downloads the prebuilt binary for your platform on first run. For a permanent anydoc command, install globally with npm install -g @firecrawl/anydoc. Run anydoc --help for all options.

Node.js

npm install @firecrawl/anydoc
import { toDocument, toMarkdown, toMarkdownBytes } from '@firecrawl/anydoc';

// From a file path:
const markdown = await toMarkdown('report.docx');

// From bytes, with the format detected from the content:
const fromBytes = await toMarkdownBytes(bytes);

// Or name it, which signature-less formats (CSV) need:
const fromCsv = await toMarkdownBytes(bytes, 'csv');

// Or stop at the document model, which also carries embedded assets:
const document = await toDocument(bytes);

Full API reference: node/README.md

Python

pip install firecrawl-anydoc
import anydoc

# From a file path:
markdown = anydoc.to_markdown("report.docx")

# From bytes, with the format detected from the content:
markdown = anydoc.to_markdown_bytes(data)

# Or name it, which signature-less formats (CSV) need:
markdown = anydoc.to_markdown_bytes(data, "csv")

# Or stop at the document model, which also carries embedded assets:
document = anydoc.to_document(data)

Full API reference: python/README.md

Rust

cargo add anydoc
// From a file path:
let markdown = anydoc::to_markdown("report.docx")?;

// From bytes, with the format detected from the content:
let markdown = anydoc::to_markdown_bytes(&bytes, None)?;

// Or name it, which signature-less formats (CSV) need:
let markdown = anydoc::to_markdown_bytes(&bytes, anydoc::Format::Csv)?;

// Or stop at the document model, which also carries embedded assets:
let document = anydoc::to_document(&bytes, None)?;

Features

Supported formats

Format Extensions
Word .doc, .docx, .docm
PowerPoint .ppt, .pps, .pot, .pptx, .pptm, .ppsx, .ppsm
Excel .xls, .xlsx, .xlsm, .xlsb
OpenDocument .odt, .ods, .odp
Rich Text Format .rtf
EPUB .epub
CSV .csv
PDF .pdf

Benchmark

anydoc is measured against six other converters on 100 real-world documents spanning fourteen formats. Scores run from 0 to 100, higher is better; speed is the median time to convert one document.

tool formats median ms docs judged score completeness structure formatting cleanliness
anydoc 14/14 4.7 94 80 88 78 77 79
libreoffice 12/14 1129.5 87 40 59 43 43 24
unstructured 8/14 572.9 58 65 76 62 52 67
markitdown 6/14 134.8 33 65 80 67 61 53
pandoc 5/14 102.1 34 57 75 57 58 39
docling 4/14 513.6 21 57 63 59 57 52
mammoth 1/14 52.5 8 70 85 68 74 55

Per format, like for like:

format anydoc libreoffice unstructured markitdown pandoc docling mammoth
doc 88 58 68 - - - -
docm 82 49 - - - - -
docx 86 53 56 72 68 68 70
epub 74 - 74 77 53 - -
odp 87 22 - - - - -
ods 82 42 - - - - -
odt 80 52 70 - 61 - -
ppt 80 25 - - - - -
pptx 76 22 - 59 - 50 -
rtf 89 58 48 - 46 - -
xls 77 40 68 64 - - -
xlsm 70 30 - - - - -
xlsx 70 31 69 55 - 51 -

How quality was scored: an LLM judge (Claude Sonnet 5) compares two tools' outputs blind against ground truth: the document's first six pages, rendered to images by LibreOffice. Each output is scored on completeness, structure, formatting, and cleanliness. Every pair is judged twice with the outputs swapped to cancel position bias, for 479 verdicts in total. Each tool's score averages its per-format scores over the formats it supports, so a corpus heavy in one format can't skew it. It also means each row averages a different set of formats (mammoth's 70 is docx alone, while anydoc's 80 spans all fourteen), so the per-format table is the fair comparison.

Speed is one warm conversion per document on a Ryzen 9 9950X3D (Windows 11, 64 GB DDR5-6400). anydoc and the Python libraries are timed with process spawn excluded; the CLI tools include it, since that is how they are used. The harness lives in bench/; the corpus is not redistributable and is not in the repo.

Best fit: pipelines that receive a mixed bag of office documents and need one consistent, structured Markdown output. In this comparison, anydoc was the only tool to cover all fourteen formats, scored highest on every judged format except EPUB, and converted documents an order of magnitude faster than the next-fastest tool.

Format detection

The format is read from the file content, using the marker its specification designates: the PDF header, the RTF open group, OLE stream names, the ZIP package mimetype and content types. CSV has no such marker, so the extension or an explicit format names it instead.

Format::from_bytes(&bytes); // Some(Format::Docx), or None when nothing matches
Format::from_extension("pptm"); // Some(Format::Pptx)
Format::from_path(Path::new("report.odt")); // Some(Format::Odt)

The same three functions exist in Node (formatFromBytes, ...) and Python (anydoc.format_from_bytes, ...).

How it works

document bytes

  ├─► format detection      → content markers, not the extension

  ├─► format parser          → one per format (doc, docx, ppt, pptx, xls,
  │                            xlsx, odt/ods/odp, rtf, epub, csv)
  │         │
  │         └─► Document     → shared model: blocks, inlines, tables,
  │                            footnotes, assets
  │               │
  │               └─► GFM serializer → Markdown

  └─► PDF → pdf-inspector    → Markdown directly

Because every format funnels through the same document model and serializer, output quirks get fixed once. A table-escaping fix for docx is automatically a table-escaping fix for rtf, odt, and everything else.

Development

cargo test
cd node && npm install && npm run build && npm test
cd python && pip install maturin && maturin develop && python -m unittest discover -s tests

A committed fixture corpus under tests/fixtures/ is snapshot-tested, tests/robustness.rs mutation-tests every fixture, and fuzz/ carries cargo-fuzz targets per format. The speed and quality benchmark lives in bench/.

Releases are tagged v<version>, which publishes the crate, the npm package, and the PyPI wheels from .github/workflows/release.yml. The version lives in three places, bumped together for a release:

License

MIT

Join libs.tech

...and unlock some superpowers

GitHub

We won't share your data with anyone else.