Docling vs Mistral OCR — Ingestion E2E Test

hermes#86 · 2026-07-17 · 5 Goldman broker PDFs (Jul-16 corpus, 50 pages) · same pipeline, one node swapped · isolation target embeddings_docling_test · all numbers live-verified from n8n execution data (execs 41212/41213 docling · 41074 Mistral prod)

FINAL VERDICT (viskares T-deep re-check, 2026-07-17): text swap NO-GO — image lane GO

NO-GO Do not swap Mistral → docling for text OCR as-is. viskares' deep re-check (5 parallel per-doc lanes + mechanical cross-scans) found docling narrative fidelity worse on 5/5 documents, unanimously — and the disqualifier is silent content loss: dropped footnotes (BHP's ENGIE water-supply footnote has zero trace in docling), dropped Reg AC sentences (3/5), truncated FINRA clauses (5/5) — all while reporting status="success". Systemic prose defects: em-dash deletion that fuses words ("settersincluding"), 75 HTML-entity leaks vs 0, 22 mid-sentence ordinal injections vs 0.
RETIRED FEAR Numbers are PARITY — ~200+ narrative-sentence figures checked, 0 disagreements. Neither engine corrupts digits in prose. (This corrects §5's earlier "indicative docling digit win": that case was a dense table row — unadjudicated without the source PDF, and tables are demoted per operator ruling.)
DECISIVE ASYMMETRY Mistral's defect (page furniture in 6–10% of lines) is proven strippable by a line filter, zero information lost. Docling's defects are unrecoverable — no regex un-drops a footnote. A pipeline can clean Mistral; it cannot un-lose ENGIE.
IMAGE LANE GO The image findings stand untouched by the prose verdict: 17 charts detected + classified at $0, captions captured, real PNGs exported (§7–§8). Recommended path: HYBRID — keep Mistral text + add a Mistral furniture-strip filter, run docling as the parallel image lane, revisit the text swap only after docling's retest gates (§9) pass, above all the content-loss root cause.
REV-2 — MISTRAL CORRUPTS TABLE DIGITS IN PROD viskares independently verified the §5 Escondida case by the document's own arithmetic: Mistral fails 4/7 consolidation columns; docling reconciles 6/6. Mistral's failure mode is the dangerous kind — silent, plausible wrong digits (0→6 misreads) in the dense broker tables Viska ingests today. Docling's prose damage is visible; Mistral's digit damage gets trusted. Neither engine is production-safe as-is; "stay on Mistral" is not a null option. New gate 7 (§9): audit prod embeddings_main for table-digit blast radius — arguably more urgent than gates 1–6. The table-demotion ruling's premise ("Mistral already emits tables") holds for structure but NOT for digit correctness — whether that reopens the demotion is the operator's call (routed via viska-pm).

1 · What was tested

Exact clone of the production ingestion workflow (hHs57Y8Iz2FjWgsD) with node #23 Extract text (Mistral OCR API) replaced by an HTTP call to docling-serve on the N5 NAS (100.100.91.47:5001, CPU standard pipeline, tables=accurate). Everything downstream — 3000/300 splitter, text-embedding-3-small, dedup, metadata — untouched. Writes isolated to the shadow table embeddings_docling_test (ViskaDB Option B ruling); hash_id kept identical to prod for 1:1 row joins. The Mistral baseline is this morning's real production run over the same 5 PDFs — same documents, same day, no re-run needed.

2 · Run evidence

CheckResultEvidence
End-to-end greenPASSexec 41212 (2m48s): docling 5/5 success → 117 chunks embedded into embeddings_docling_test, 0 error routes
Idempotent re-runPASSexec 41213 (1.4s): dedup read the shadow table back (19/16/48/16/18 = exactly the run-1 chunk counts), all 5 skipped, 0 API calls, 0 writes
Prod corpus untouchedPASSzero embeddings_main / ledger writes in both runs (ledger writers held disabled by design)
Dense table as real markdownPASSBHP/S32 doc: 381 markdown table rows incl. multi-column synergy & financials tables
Live Dropbox pathCONFIRMED/Current/2026/July/Jul 16/… (65 PDFs found; the /Reports/… doc is stale)

3 · Speed — Mistral API is ~4.7× faster than docling on N5 CPU

Per-document OCR time, same PDFs. Mistral = hosted API (network + parse); docling = local CPU parse on the N5. Totals: 30.5s vs 143.9s for the 5-doc set (50 pages).

UMT · 7 pg
6.3s
22.8s
ECB Chatterbox · 7 pg
3.8s
8.9s
BHP/S32 Copper · 19 pg
11.3s
90.7s
LATAM Today · 10 pg
4.2s
6.0s
Kawasaki Kisen · 7 pg
4.9s
14.6s
Mistral OCR (API) Docling (N5 CPU)

Context that softens this: docling is self-hosted, so there is no 429 throttle, no 5s pacing requirement, and no per-page bill — the daily batch is ~65 docs and even at docling speed fits comfortably in the ingestion window. Speed only becomes a blocker for backfills, where the async endpoint or parallel workers would be needed.

4 · Structure — both engines emit markdown tables (the spec's premise was wrong)

DocumentMistral table rowsDocling table rowsRead
UMT7060comparable
ECB Chatterbox4039comparable
BHP/S32 Copper447381Mistral splits per page (repeated headers); docling merges some rows — see §5
LATAM Today00narrative doc, no tables — agreement
Kawasaki Kisen31–4531comparable

The hermes#86 hypothesis said docling would win because "Mistral flattens tables to text." Live prod output disproves that — the Mistral OCR node already produces markdown headings, bold, bullets and pipe tables. Any cutover case must rest on fidelity, cost and independence, not structure.

5 · Fidelity — the decisive case study (Escondida production row)

VERIFIED CORRECT — viskares REV-2 §10 [live-verified]: this case study was briefly marked superseded, then independently adjudicated by viskares using the document's own arithmetic (Escondida + South America ex-Escondida + South Australia must equal the Copper-Consolidated row, present in both engines): Mistral fails 4 of 7 columns; docling reconciles 6 of 6. Three independent lines converge — internal arithmetic, the guidance band (1,000–1,100: docling inside, Mistral ~60% above), and real-world sanity (Escondida runs ~1.0–1.3 Mt/yr; 1,663kt would be an all-time record). Mistral's pattern: 0→6 digit misreads (1,063→1,663; 1,033→1,633) plus 304→354, 1,956→1,996. These are silent, plausible wrong numbers — in production today. Narrative-number parity (200+ checked, 0 disagreements) still holds; the corruption is specific to dense tables.

Same table, same row, both engines. The document's own guidance row sits directly beneath it and provides an internal consistency check:

Mistral (prod, exec 41074) clean format, suspect digits
| Excondida | kt | 1,125 | 1,305 | 1,261 | 1,663 | 1,633 | 974 | 969 |
| Guidance  |    |       |       |       | 1,000-1,100 |   | 900-1,000 | |
Docling (test, exec 41212) messy cells, consistent digits
| Pr oduct i o n E sco n d i da | k t | 1,1 25 | 1, 305 | 1, 26 1 | 1, 063 | 1, 033 | 9 74 | 969 |
Why this matters: Mistral reads 1,663 / 1,633 kt where docling reads 1,063 / 1,033 kt. The document's own guidance row for that column says 1,000–1,100 — docling's values sit inside it, Mistral's are 50% above it. This strongly suggests Mistral committed 0→6 digit misreads (plus the entity typo "Excondida"). A wrong digit is the worst failure class we have: it survives chunking, embedding, retrieval, and lands in an analyst-facing answer looking authoritative. Docling's spacing noise ("1, 063") is ugly but the digits are recoverable — and an LLM reading the chunk will usually normalize it. Final adjudication still requires the source PDF (viskares' leg).

Defect classes observed (5-doc sample)

EngineDefectSeverity for RAGExample
MistralDigit misreads in dense numeric tablesCRITICAL silent, survives retrieval1,663 / 1,633 vs guidance 1,000–1,100
MistralEntity name corruptionMODERATE hurts entity retrievalExcondida (Escondida)
DoclingIntra-cell character spacingMODERATE noisy but recoverablePr oduct i o n E sco n d i da, 1,1 25
DoclingAdjacent row merges in dense appendix tablesSERIOUS two rows' numbers in one cellBrazil Alumina T ota l Al u min a … 1,286 5 , 063
Doclingfi-ligature artifactMINOR 1 occurrence in 264KB昀 nancial
DoclingWhole-doc output → page numbers lostMODERATE weakens citationsall chunks carry page_number=1

Numeric-token overlap on the BHP tables (indicative, not adjudicated): 630 unique tokens shared, 257 Mistral-only, 101 docling-only — most deltas are the spacing/merge artifacts above, some are real digit disagreements. Only the source PDF can score them.

6 · So — is it an improvement?

AxisWinnerBasis
Digit fidelity (load-bearing)Docling, indicativelyGuidance-consistency case study; awaits PDF ground truth
Throttle / availabilityDoclingSelf-hosted: no 429s, no 5s pacing, no repeat of the Jun-10 Mistral-401 incident (321 files silently skipped)
Marginal costDocling$0/page on owned hardware vs metered OCR API
LatencyMistral~4.7× faster; docling fine for the daily window, needs async/parallel for backfills
Cell hygiene in dense tablesMistralNo spacing corruption or row merges
Table structure (markdown)TieBoth emit markdown tables — premise correction
Page attributionMistralPer-page items; docling whole-doc (fixable via json_content or per-page split)

Reading: if the PDF spot-check confirms the Escondida pattern — docling right, Mistral wrong on digits — the cutover case is strong despite the speed and hygiene regressions, because digit errors are silent and analyst-facing while docling's defects are visible and mostly mechanical (spacing normalization and a row-merge guard are cheap post-processing; page attribution is recoverable from json_content). If the spot-check instead shows both engines erring, the answer becomes a post-processing bake-off rather than a swap.

7 · Image probe (step 2, operator-requested) — narrative + image retrieval is REAL

Follow-up run on the same 5 PDFs with do_picture_classification=true (execs 41216 sync + 41218 async, 2026-07-17 afternoon). Operator ruling applied: speed deprioritized, cost + consistency + image capability primary.

DocumentPagesPicturesCharts (by type)Noise filteredParse + classify
UMT721 line1 logo39.0s
ECB Chatterbox711 logo16.6s
BHP/S32 Copper19158 line + 2 pie + 2 bar (+ 1 geographic map, 1 engineering drawing)1 logo107.8s (async)
LATAM Today1032 line1 logo35.6s (async)
Kawasaki Kisen732 line1 logo22.0s
Total502417 charts (+2 map/drawing)5 logos
What this proves: docling detects every figure with page + bounding box, and classifies it — pie charts at 0.9999 confidence, bar charts at 0.996, the Goldman banner logo correctly separated on every single document. 17 retrievable charts in 5 narrative PDFs, at $0 marginal, on the same call the OCR already makes. The taxonomy is richer than expected (geographical_map, engineering_drawing — the BHP mine map and site diagram, both genuinely useful to a client). Figure descriptions do not need the N5 GPU: picture_description_api accepts any OpenAI-compatible endpoint, so a cheap mini-model can caption charts for embedding.

Production constraint found (and solved in-probe)

The sync endpoint caps at DOCLING_SERVE_MAX_SYNC_WAIT=120s and a single slow document starves everything queued behind it (a 6-second doc 504'd waiting behind the 19-page BHP note). The async endpoint (submit → poll → result) ran the same docs cleanly through n8n — production ingestion must use it. With the multi-hour window this is a non-issue operationally.

8 · Retrieval example — real extracted assets from ONE document

Everything below was extracted by docling from Global Metals & Mining: BHP/S32 Copper.pdf (Goldman, Jul 16) in this test — the images are the actual PNGs docling cropped out (exec 41224, scale 2, with page + class + caption), the quotes are verbatim from its markdown output. Only the question framing is illustrative; the production join is mechanical: text chunk hit → same hash_id + page → attached figures → source-doc link.

Client asks Mímir: "What could BHP unlock from the Spence–Sierra Gorda partnership?"

Answer (synthesized from the retrieved narrative)

Goldman assesses that BHP's Spence mine and the Sierra Gorda JV — 20 km apart in Northern Chile — could unlock US$2–4bn in total value under their new MoU: roughly US$1.0–1.7bn for Spence and US$1.1–1.9bn for Sierra Gorda, via shared cathode and milling capacity, water, power, tailings and procurement. Goldman values Spence at US$8.4bn (~3% of BHP NAV) and Sierra Gorda at US$10.4bn, and reads the deal as a template for further asset-level JVs in the Americas.

Supporting quotes (verbatim, as-extracted)

  • "We assessed potential synergies between the two mines and concluded that it could unlock US$2-4bn in total value through collaboration on spare cathode and sulphide milling and mining capacity, as well as water, power, tailings, and procurement…" (p.3)
  • "Sierra Gorda (SG) has an oxide ore stockpile containing ~320kt of copper that could be trucked and processed at the Spence dynamic heap leach and cathode plant, which has ~130ktpa of spare cathode capacity… This could generate ~US$200mn p.a. of EBITDA on our estimates." (p.3)
  • "We value Spence at US$8.4bn (~3% of BHP's NAV) and Sierra Gorda at US$10.4bn… The potential ~US$1.3bn (mid-point) for Spence would represent ~US$0.25/sh… We note these synergies represent ~10-20% of our NPV of the assets." (p.6)

Extraction hygiene note: two stray characters in the raw quotes ("Sierra n Gorda", "stockpile o containing") were removed in the second quote and left intact in the third's source — this is the docling narrative noise level on this doc: readable, minor, mechanically cleanable.

Retrieved figures (actual docling-extracted PNGs)

Satellite map: Sierra Gorda SCM and Spence Mine, 20km apart
geographical_map · p.4 · conf 0.91 · docling-captured caption: "Exhibit 1: BHP's Spence copper mine and the Sierra Gorda (KGHM 55%/South32 45%) in Northern Chile are just 20km apart"
BHP NAV breakdown pie charts
pie_chart · p.9 · conf 0.9999 · BHP NAV mix (Copper US$113.1bn / 43%, Escondida 16%) — the denominator behind the "~0.5% NAV uplift" quote
BHP rating and price target history
line_chart · p.15 · conf 0.95 · GS rating & price-target history (BHPB.L)

Source document for client inspection: /Current/2026/July/Jul 16/Goldman/Global Metals & Mining_ BHP_S32 Copper.pdf — served on demand via Dropbox temporary link. Join key: hash_id e0c121c0…987e68 (identical across text chunks, figures, and the prod Mistral rows).

9 · What happens next — viskares' six retest gates + the hybrid build

viskares verdict of record: /tmp/war-room/viska-main/artifacts/docling-e2e/NARRATIVE-VERDICT.md. A future docling text retest requires all six gates; #6 is the blocker — silent content loss must be root-caused before docling can be trusted on documents nobody hand-diffs.

#FixOwnerRetest gate
1Decode HTML entities (&) in md outputViskaN8N post-process0 entity hits
2Root-cause em-dash deletion + word fusiondocling config / proteusem-dash count matches Mistral; 0 fused words
3Root-cause mid-sentence ordinal injection + stray glyphsdocling config0 mid-sentence ordinals
4Page attribution via per-page json_contentViskaN8Npage_number spans real range
5Preserve heading hierarchy (h1–h4, not flat h2)docling configh1–h4 present
6Explain dropped footnote / Reg AC / FINRA content lossdocling config5/5 parity on probe sentences — disqualifying until understood
7REV-2: root-cause Mistral's silent table-digit corruption; audit prod embeddings_main for blast radius across all ingested broker tables — wrong numbers are in the RAG corpus answering questions todayViskaN8N / ViskaDB / operatorcorruption rate bounded; affected rows identified — arguably more urgent than gates 1–6

Actionable now (independent of the text NO-GO)

StepOwner
Mistral furniture-strip filter in prod ingestion (viskares proved the line filter restores severed sentences; 6–10% of lines are furniture polluting chunks today)ViskaN8N
Docling image lane (parallel to Mistral text): async extract + classify → chart PNGs to storage → captions embedded → narrative + chart + doc-link answers (§8 demo; lands backlog 006/007)ViskaN8N + ViskaDB (bucket)
Shadow-table teardown (TRUNCATE/DROP embeddings_docling_test) once evidence no longer neededViskaDB
Route gates #2/#3/#5/#6 to the docling-serve owner (config/version investigation on N5)proteus via hermes#86

Test workflow XUkhbVILviu5deZC deactivated after runs. Artifacts: /tmp/war-room/viska-main/artifacts/docling-e2e/ (RESULTS.md, 5× docling markdown, 5× Mistral text, diff-join-keys.tsv). Evidence tags: all execution facts [live-verified] via n8n REST; the fidelity call is [indicative] pending PDF ground truth.