psikosen/t-rsi-ocr

🤗 Hugging Face 来源image-to-textmit1M 参数1 MBsafetensors✓ 4 个校验和今天更新
需要做种者 →

T-RSI Fast Reflex OCR

A 1.58-bit ternary BitNet text-line recognizer in 494 KB, with a dependency-free C engine

[!WARNING] Experimental research release. Every number below was measured on one Apple Silicon machine against ground truth this project generated. It has not been independently replicated. Treat the figures as reproducible-by-you, not as verified.

Every weight is -1, 0 or +1, packed four to a byte. The model reads a 32-pixel-tall text strip and emits characters through CTC, decoded with a prefix beam search over a 27 KB character language model.

It ties or beats Tesseract 5 on character accuracy in every scenario measured, reads a line 8x faster in 5.6x less memory, and still returns fewer perfect lines. Full numbers and the reasons are below, including where it loses.

That speed figure is ARM-specific. It was measured on an Apple M4 Max, where the C engine uses NEON and the ARMv8.2 dot-product instruction. Those carry the engine: without any SIMD it runs at 144 ms/line, which is Tesseract's speed. See Hardware dependence before quoting a number.

The model weights are unchanged in this release. v4 of the engine is a decode-time and implementation release: 3.9x faster per line and 4.8x more throughput with threading, at identical output. The recogniser is exactly as accurate as it was; it stopped wasting its own output and stopped wasting cycles. What changed is recorded under Engine v4.

The exact-line gap used to be 25.7 points and is now 3.6 to 6.6 depending on the sample. What closed it is multi-stretch fusion, described below: decode the same line at two stretches, align the frames, average the log-probabilities, and decode once. It replaced the hypothesis voting an earlier version of this card recommended, and is both more accurate and about 6x faster than it.


Measured results

Reproduce with python3 benchmark_vs_tesseract.py. Ground truth needs no hand labelling: a digital PDF already contains its own exact text, so rasterizing a page and cropping each line at the bbox the PDF reports gives pixels whose correct answer is known. Training and test documents are split by file and deduplicated by content fingerprint.

Character accuracy

scenario T-RSI Tesseract 5
fonts, clean 99.50% 98.09% +1.41
fonts, degraded 99.27% 96.97% +2.30
documents, clean 98.18% 97.86% +0.32
documents, degraded 0.5 97.83% 97.46% +0.37
documents, degraded 1.0 95.57% 93.57% +2.00

On the larger 3,000-line corpus in the deep benchmark below, clean documents come out a tie rather than a win (97.72% against 97.62%). Two samples, two answers, both reported.

Exact-line match — where it loses

scenario T-RSI Tesseract 5
fonts, clean 93.7% 75.7%
fonts, degraded 92.3% 70.3%
documents, clean 84.3% 87.9%

Tesseract's errors cluster onto a few hard lines and leave the rest clean. This model's scatter — single characters across many otherwise-good lines — and one stray character ruins a line as thoroughly as ten. That is the documented weakness of CTC, which treats every timestep as conditionally independent given the image and so cannot notice that the letters it just emitted do not form a word.

Multi-stretch fusion closes most of that gap, and is the default. Decode the same line at more than one horizontal stretch, align the frames with DTW, average the LOG-PROBABILITIES, and decode once. The stretches fail on different characters, so averaging lets them cancel.

Combining the numbers rather than the decoded strings is the point. An earlier version of this card recommended voting on decoded strings (ROVER); measured head to head on the same lines, aligning and averaging the log-probabilities beats it by 1.8 points of exact-line match, because a decoded string has already thrown away how sure the decoder was. Soft decoding beating hard is a Shannon-era result, and confusion-network combination beats ROVER in speech for the same reason.

C engine, 2,000 held-out lines, latency on an M4 Max:

decoding char accuracy WER exact line ms/line
--fuse=single 97.16% 0.0841 71.70% 7.5
--fuse=auto (pair, only when unsure) 97.39% 0.0742 76.50% 11.9
--fuse=pair (default) 97.39% 0.0739 76.25% 14.2
--fuse=autofull (all five, only when unsure) 97.53% 0.0708 79.00% 23.3
--fuse=full (all five, always) 97.53% 0.0710 79.00% 33.3
4 strings + ROVER vote 97.46% 0.0733 78.05% —
Tesseract 5 97.51% 0.0503 84.60% ~106

Latency is the minimum of seven interleaved runs, not the median. Benchmark interference only ever adds time -- nothing runs faster than the hardware allows -- so on a contended machine the minimum is the robust estimate and every higher sample is measurement error. This is not pedantry: a median over one noisy run reported autofull at 37.6 ms and full at 34.8, reversing their order. The ratio between them held at 0.70 across a 2.3x swing in machine load.

autofull reaches full's accuracy for about 30% less work. Identical exact-line match and character accuracy, because on the lines it skips, pass one was already right. Its risk is the mirror of its gain: where auto degrades toward pair if the gate stops firing, autofull degrades toward single, and that is 7.3 points rather than 4.5. Higher ceiling, further to fall.

auto decodes once, scores its own answer with the CTC forward algorithm, and pays for the second stretch only when that score is low. It is cheaper and slightly more accurate than fusing everything, because fusion occasionally hurts a line the first pass already had right. It is opt-in rather than the default because its threshold is calibrated on clean digital PDFs: if confidence drifts high on very different input the gate stops firing, and that accuracy loss is silent. Re-measure before trusting it on scans.

Fusion costs one extra forward pass, no extra bytes, and no retraining. Alignment has to be DTW: assuming a stretch scales time uniformly and resampling linearly gives up 1.5 points, because the model does not fire on a uniformly scaled schedule.

Which mode to use

Spend on list depth before you spend on fusion depth. Extra candidates are free -- the beam has already ranked them -- while every extra stretch costs a whole forward pass. The two buy the same thing and one of them is 2.4x cheaper:

exact line ms/line
pair, top-5 84.65% 14.04
full, top-3 84.30% 33.10
pair, top-8 85.40% 14.04
full, top-5 85.40% 33.10

So full is the best accuracy, not the best mode. It buys 2.75 exact-line points over pair for 2.4x the compute, and if a shortlist is acceptable to whatever consumes the output, a longer list from pair gets there for nothing.

A rough decision:

  • One answer, throughput matters -> pair. The default, and the balanced point: 8x Tesseract's speed at a tie on character accuracy.
  • One answer, accuracy matters more than cost -> autofull on input resembling clean digital PDFs, full otherwise. Same accuracy; autofull is ~30% cheaper but inherits the gate's silent failure mode, and full has no threshold at all. For an archive that must not quietly degrade, take full and the predictability.
  • A shortlist is acceptable -> pair with --nbest=5. Passes Tesseract's single answer, at the default's cost.
  • Large batches of clean digital PDFs -> auto, which is cheaper and marginally more accurate than pair. Read its caveat above first.
  • Throughput above all -> single with --threads=4. 71.75% exact, but roughly 2 ms/line.

auto dominating pair on both axes is not a rounding artefact: fusion occasionally damages a line the first pass already had right, so not fusing the confident ones is a small accuracy gain as well as a saving. It is still not the default, because its threshold is a constant measured on one kind of document and the failure is silent.

By content type (real documents)

content T-RSI Tesseract 5
prose 99.53% 99.23%
symbol-heavy 97.41% 96.71%
numeric / amounts 99.09% 98.71%
alphanumeric IDs 83.71% 84.67%
ALL CAPS headers 99.60% 100.00%

Alphanumeric identifiers are the weakest content for both engines. Random codes carry no language structure, so no language model can rescue a misread glyph.


Deep benchmark vs Tesseract 5

3,000 held-out lines from 8 unseen documents, three degradation levels. Run it yourself with python3 deep_benchmark.py.

Two corpora appear on this card and their numbers differ. This section uses 3,000 lines from 8 documents; the decoding and n-best tables use 2,000 lines from 6, which is the set the engine was tuned against. Exact-line match reads 78.9% on the first and 76.25% on the second, for the same binary and flags — that spread is corpus, not configuration. Each table states which it used.

metric T-RSI (C) Tesseract 5
Character accuracy, clean 97.72% 97.62% tie
Character accuracy, degraded 95.33% 93.83% win
Word error rate 0.0658 0.0545 loss
Exact-line match 78.9% 85.5% loss
(2,000-line corpus, for reference) 76.25% 84.60% loss
Peak RAM (50-line batch) 6.3 MB 35.3 MB win 5.6x
Cold start 19 ms 94 ms win 5.0x
Throughput (ARM, NEON+SDOT) 13.3 ms/line 106.1 ms/line win 8.0x
Install size 494 KB 48.9 MB win 101x

5 wins, 2 losses, 1 tie.

Tesseract's own timing is noisy and the ratio moves with it. Three runs over the same 3,000 lines measured it at 147.8, 103.9 and 106.1 ms/line, while this engine held steady. 8.0x is the like-for-like figure from the run above; an earlier version of this card said 8.8x, which came from Tesseract's slowest run.

Peak RAM is batch-dependent and the 6.3 MB above is a 50-line figure. It grows sub-linearly with batch length in every build measured -- 8.2 MB at 100 lines, 15.7 at 1,000, 24.0 at 2,000, 37.8 at 4,000 -- because per-line buffers are sized to each line's frame count. Quote it with the batch size attached.

Per document -- genuinely mixed

4 wins, 2 losses and 2 ties on character accuracy across the 8 documents. One file swings it either way; this is not a clean sweep.

Content type

content n T-RSI Tesseract
prose 1967 98.44% 98.28%
symbol-heavy 784 97.29% 96.67%
alphanumeric IDs 134 92.16% 92.84%
numeric / amounts 69 95.30% 98.63%
ALL CAPS headers 46 94.14% 98.37%

Numbers and capitals are the weak spots, by 3.3 and 4.2 points. If your documents are mostly figures, this is the wrong tool.

Line length

length n frames/char T-RSI Tesseract
< 30 chars 561 1.39 92.40% 93.63%
30-60 767 1.25 97.88% 97.54%
60-90 506 1.25 99.49% 99.03%
90-120 1158 1.25 99.43% 98.99%

It wins on long lines and loses on short ones -- and short lines have the best frame budget, so this is not the deletion problem that stretching solves.

This is the biggest unexplained loss in the model, and it survived a dedicated hunt. Tried and measured on short lines, none of which helped: turning the character LM off (it costs 2.7 points, not gains), priming the LM with a leading space, raising or lowering the insertion bonus, a wider beam, and padding the line with white margin so the first glyph is not at frame zero. The only thing that helps is more stretch, which fusion already does per line. Short lines are headers, captions and table cells set in larger type; why that hurts is not known.


C engine

csrc/trsi_ocr.c -- one file, no dependencies beyond libm. No BLAS, no image library; input is binary PGM parsed in twenty lines.

cd csrc && make
./trsi_ocr ../model_v4.trsi line.pgm          # fuses 2 stretches, the default
./trsi_ocr --fuse=full ../model_v4.trsi line.pgm    # 5 stretches, more accurate
./trsi_ocr --fuse=single ../model_v4.trsi line.pgm  # one decode, fastest
./trsi_ocr --self-check                       # verify the SIMD paths

It is a port of engine.py, not a generic BitNet runtime. bitnet.cpp (arXiv:2410.16144) provides ternary matmul kernels for LLMs but has no patch tokenizer, no CTC, no beam search and no character LM, which is most of what this model is.

Two measured notes from building it:

  • The first version ran at 364 ms/line because it unpacked 2-bit codes inside the inner loop. Unpacking once at load costs 1.25 MB resident and made it 5.7x faster. At 6 MB, memory was no longer the constraint.
  • NEON SIMD took it from 63.7 to 23.7 ms/line, bit-identical on all 300 test lines at the time.
  • Profiling the fused engine put 97% of the time in two functions: the ternary matmul (50.7%) and attention (46.5%). The DTW alignment that fusion is named after was 0.6% -- fusion was never slow because of the fusion, only because it runs the transformer twice.
  • Rewriting those two took it from 27.5 to 9.2 ms/line unfused and 52.0 to 17.5 ms fused, so the fused path now costs less than the unfused one used to. The matmul moved from float FMA to SDOT (int8 dot product, 16 multiply-accumulates per instruction) -- the old code widened ternary weights to float, doing multiplication on a format that exists to avoid it. Attention gained SIMD, a loop interchange that turned a stride-192 read of v into 32 contiguous floats, and a 4-wide polynomial exp in place of ~1.3 million libm expf calls per line.
  • That last change gives up bit-identity by design, so it is bounded instead: ./trsi_ocr --self-check asserts the fast exp within 1e-5 relative of libm (measured 4.97e-6) and monotone, and the SDOT product within 1e-5 of a scalar reference (1.02e-6). End to end, exact-line match on 2,000 held-out lines moved by at most 0.30 points. Scalar fallbacks are kept for CPUs without the instructions.
  • Measured and rejected: -O3, -funroll-loops and -march=native are all within noise of plain -O2 here (17.4-17.9 ms/line). Threading was not attempted -- a line recognizer is normally called on many lines, so the parallelism belongs in batching one level up.

model_v4.trsi (410 KB) is a flat binary holding the packed weights, the character LM and the decode parameters, in the exact order the C reads them -- no JSON parser, no name lookup.

Threading

A batch is embarrassingly parallel: nothing is shared but the read-only model. --threads=N recognises N lines at once, --threads=0 uses every core. Threads pull the next line off a shared counter rather than taking a fixed slice, because line widths vary several-fold and attention is quadratic in width -- a static split leaves cores idle behind the longest chunk. Output is written to each line's own slot and printed in input order, so it does not depend on thread count or scheduling; 16 threads is byte-identical to 1 on all 2,000 test lines.

threads ms/line lines/s peak RSS MB per line/s
1 14.05 71.1 23.9 MB 0.336
2 6.98 143.3 27.0 MB 0.188
4 4.10 243.9 32.5 MB 0.133
8 3.07 325.7 44.3 MB 0.136
16 2.71 369.7 60.8 MB 0.164

Four threads is the cost optimum, not sixteen. Past four you pay ~50% more memory for ~30% more throughput, and the last column -- memory per unit of throughput, which is what an instance bill tracks -- turns back upward. Scaling is 3.4x on 4 cores and 5.2x on 16, not linear: the cores are not uniform and the ternary matmul touches memory even though the weights are small.

Engine v4

The model weights did not change. Everything below is decode-time or implementation, and every accepted change produced byte-identical output or better accuracy.

change effect
Ternary matmul: float FMA -> int8 SDOT 27.5 -> 18.2 ms/line
Attention: NEON dot, loop interchange, polynomial exp 18.2 -> 9.2
4-way output unroll (4 independent SDOT chains) −11.8%, bit-identical
2 x 4 register blocking (each weight load used twice) −4.9%, bit-identical
Vectorised the remaining exp loops −1.7%
Trimmed dead lm_head rows −1.1%, identical
Per-thread scratch arena −8% peak RSS, identical
Threaded batch 5.2x throughput, deterministic

Profiling note worth repeating: an instrumented build reported softmax at 66% of runtime and totalled 33 ms/line for a binary that runs at 15. The timing calls cost more than the code they measured. Ablation -- delete a phase, diff the wall clock -- gave the real ranking: matmul 45%, attention 25%, the rest 30%.

Hardware dependence

Every throughput figure on this card is from one machine, an Apple M4 Max, and the speed is mostly the SIMD, not the C. The ternary matmul and attention are 97% of the runtime, so how well they vectorise decides everything.

Three build paths, same source, measured on that machine by disabling the higher ones. Fused two-stretch decode, 100 held-out lines, median of three runs:

build ms/line selected on
NEON + SDOT 16.6 ARMv8.2+ with dot product: Apple Silicon, Graviton3, recent Cortex-A
NEON, no SDOT 23.1 any ARMv8: Cortex-A53/A72, Raspberry Pi 3/4
no hand SIMD, compiler still vectorises 20.7 -DTRSI_NO_SIMD
no SIMD at all 144.1 -DTRSI_NO_SIMD -fno-vectorize -fno-slp-vectorize

Two things worth reading off that table.

The bottom row is the honest floor. With nothing vectorised the engine runs at 144 ms/line, against Tesseract's ~106 — the speed advantage is an argument about SIMD, not about being small. The memory and install-size wins are unaffected.

Hand-written intrinsics are worth less than they look (16.6 against 20.7), because clang at -O2 on arm64 already vectorises the plain loops. The intrinsics buy the SDOT path, which the compiler will not choose on its own.

x86: it works, and it is leaving the main optimisation on the table.

There is no x86 machine in this evaluation, so there is no native x86 timing on this card and none should be inferred. What could be established on ARM:

  • An -arch x86_64 build compiles, passes --self-check, and decodes 100 real lines with one line differing from the ARM build -- the same float-noise level that separates any two paths here. It is functionally equivalent.
  • clang auto-vectorises the ternary matmul on x86, but as floats: mulps at the x86-64 baseline, vmulps with -mavx2. It emits no integer dot-product instruction. So x86 runs the arithmetic this engine used before the SDOT rewrite -- widening ternary weights to float to multiply them.
  • x86's direct equivalent to SDOT is vpdpbusd (AVX-VNNI on Alder Lake and later, AVX512-VNNI on Ice Lake and Zen 4). The compiler here accepts it, so the instruction is reachable; it is simply not implemented in this source.

That is the concrete gap, and it is not written blind on purpose: a VNNI binary does not execute under Rosetta on this machine, so an x86 integer path could not be checked for correctness here, never mind timed. It needs real x86 hardware.

Rosetta timings are not reported. Running the x86 binary under translation measures the translator, not an x86 CPU.

The paths agree on text, not on bits. Across 100 lines the non-vectorised build produced identical output to NEON+SDOT, and the auto-vectorised one differed on a single line -- float reassociation, the same order of noise that separates the C and Python engines. --self-check prints which path a binary was compiled with.


Size, memory and speed

T-RSI (C) T-RSI (Python) Tesseract 5
Packed ternary weights 312.0 KB 312.0 KB —
Float32 tensors (LayerNorm, biases, scales) 27.4 KB 27.4 KB —
Character 3-gram LM 27.4 KB 27.4 KB —
Engine 50 KB binary 26.1 KB code —
Install total 494 KB 393 KB + numpy 48.9 MB
Peak RSS, 50-line batch 6.3 MB 55.6 MB 35.3 MB
Peak RSS, 2,000-line batch 24.0 MB — 35.3 MB
Cold start to first result 19 ms — 94 ms
Steady-state per line, 1 thread 13.3 ms 55 ms 106 ms
Steady-state per line, 4 threads 4.10 ms — 106 ms

101x smaller on disk, 8x faster per line single-threaded (26x on 4 threads), 5x faster to cold start, and 5.6x less memory — in the C engine, on ARM with NEON and SDOT.

Two caveats the numbers above need. The speed does not hold without SIMD: the same code runs at 144 ms/line unvectorised, level with Tesseract — see Hardware dependence. And peak memory grows with batch length in every build measured, so the 50-line figure is not the number a real batch will show.

In Python it uses more memory than Tesseract, 55.6 MB against 35.3 MB. The model is a small part of that: numpy and pillow cost 32.1 MB before a single weight loads. An earlier version of this card said a C port "would remove that, but no such port has been measured, so it is not a claim." It is now measured, and it is the row above.

Position encoding is fixed sinusoidal, computed at load and stored as zero bytes. A learned embedding at 1024×192 float32 would cost 768 KB — more than twice all the ternary weights combined.


Architecture

Layers 4
Hidden size 192
Attention heads 6 (head dim 32)
FFN intermediate 384
Patch grid 32 × 8, one token per 8-pixel column
Ternary weights 1,277,952
Vocabulary 256 slots, 95 printable ASCII
Weight quantization ternary {-1, 0, +1}, STE, per-tensor `mean
Activation quantization 8-bit absmax, dequantized
Decoder CTC prefix beam search, width 4
Decode-time fusion 2 stretches (1.2x, 1.3x), DTW-aligned log-probabilities

Inference is engine.py, NumPy only, no framework. Weights are ternary, so an integer SIMD or microcontroller port needs only additions and subtractions. On a CPU the same arithmetic runs through BLAS, which is not literally multiplication-free — the property belongs to the weights, not to this runtime.

engine.py is checked against the PyTorch training module on shared weights and must agree to float32 precision. That test exists because the two once silently computed different functions.


Quickstart

from PIL import Image
from engine import FastReflexOCREngine

engine = FastReflexOCREngine.from_pretrained(
    "psikosen/t-rsi-ocr",
    lm_name="char_lm_3gram.npz",   # omit for greedy CTC
    beam_width=4,
)

text, latency_us, ram_kb = engine.recognize(Image.open("line.png"))
print(f"{text!r}  ({latency_us:.0f} us)")

Two-stretch fusion is on by default. To trade accuracy against latency:

import fusion

engine.fuse = fusion.FULL     # 5 stretches: +2.75 exact-line points, 2.4x slower
engine.fuse = None            # single decode: 2x faster, -4.55 points

When a line must come back verbatim, ask for candidates instead of one answer. The beam already ranks its hypotheses, so the list is free, and the correct string is in the top 8 for 41% of the lines the top answer gets wrong:

# C engine: three ranked candidates per line, tab separated
./trsi_ocr --fuse=pair --nbest=3 --beam=8 model_v4.trsi line.pgm
# Python engine
from beam_decoder import prefix_beam_search

candidates = prefix_beam_search(
    logits, engine.blank_id, {i: t for i, t in engine.id_to_token.items() if len(t) == 1},
    lm=engine.lm, beam_width=8, n_best=3,
)

All figures below come from the same binary, the same 2,000 held-out lines and the same flags as the exact-line numbers elsewhere on this card, at beam width 8. Reproduce with python3 eval_nbest.py --n 2000 --beam 8.

list --fuse=single --fuse=pair --fuse=full
top-1 71.75% 76.25% 79.00%
top-2 79.15% 82.00% 83.20%
top-3 81.00% 83.55% 84.30%
top-5 82.90% 84.65% 85.40%
top-8 83.70% 85.40% 85.85%
Tesseract 5 (top-1) 84.60%

At top-5 the default configuration passes Tesseract's single answer, and top-8 clears it by 0.8 points. Of the 475 lines top-1 gets wrong, 38.5% have the correct string somewhere in the list.

That is not a better recognizer, it is a different contract — useful when a downstream parser knows the shape it expects, or when a human verifies the line against the original anyway and three candidates are more honest than one confident mistake.

Two things this table must not be read as. --nbest never widens the beam: asking for more candidates returns more of what the search already found, and top-1 above is the same answer the engine ships. Widening the beam with --beam is a different setting and moves top-1. And Tesseract's column is a top-1 figure — comparing our top-5 against it is only fair for a consumer that can act on a shortlist.


Two things that matter more than the model

1. Most documents need no OCR at all. DOCX files and digital PDFs already contain their exact text. Reading it takes 4.8 ms per page against Tesseract's 2,038 ms for the same page — 425× faster and exact rather than 97% accurate. Rasterizing a text layer and guessing it back loses table structure and turns 1999.00 into 999.00. Check for a text layer before reaching for any recognizer.

2. This is a line recognizer, not a page reader. It needs pre-cropped lines. Tesseract does layout analysis and recognition in one pass; this does not, and no line detector ships with it. A per-line speed comparison assumes the crops already exist.


Training

python3 train_v2.py --epochs 70 --pool 50000 --real 25000 \
    --layers 4 --hidden 192 --heads 6 --guide-weight 0.5 --out model_v4

49,022 synthetic lines across 111 screened system fonts, mixed with 19,816 real lines cropped from PDFs. Both matter: a synthetic-only model scored 81% on held-out fonts and 16% on real documents — it had learned the renderer.

Augmentation covers blur, sensor speckle, non-uniform lighting, ±5° skew, exposure drift, and horizontal scale jitter (1.0–1.4×). That last one is what makes input_stretch work at inference: CTC cannot emit more characters than it has frames, real lines average 1.29 frames per character, and stretching to 1.69 is worth +14.9 points of exact-line match. Applied to a model trained without the jitter, the same trick is a net loss.

Training uses an attention decoder as an auxiliary teacher (GTC, arXiv:2002.01276), discarded at export — it costs zero bytes and zero milliseconds at inference.


Limitations

  • Line recognizer only. No page layout, no line segmentation, no table structure.
  • 95 printable ASCII characters. No accents, no CJK, no mathematical notation, no typographic quotes.
  • Exact-line match still trails Tesseract, 78.9% against 85.5% on the 3,000-line corpus with fusion on. It was 74.6%. Ask for an n-best list if a line must come back verbatim.
  • Short lines are the weakest bucket and the least explained. Under 30 characters: 92.40% against Tesseract's 93.63%. They have the best CTC frame budget in the corpus and still the worst accuracy, so the usual explanation does not apply. Six fixes were measured and none helped.
  • The Python engine uses more RAM than Tesseract (55.6 MB against 35.3 MB), because of numpy and pillow rather than the model. The C engine uses 6.3 MB on a 50-line batch and 24 MB on a 2,000-line one.
  • Alphanumeric identifiers are weak (83.7%), as they are for every engine measured.
  • Fusion runs the model twice. All headline accuracy assumes it is on. --fuse=single is roughly 2x faster and gives up 4.55 exact-line points.
  • The speed advantage is SIMD, and it is measured on one ARM machine. With no vectorisation the engine runs at 144 ms/line against Tesseract's ~106 — a loss, not an 8x win.
  • Peak memory grows with batch length, sub-linearly, in every build measured. Any single figure needs its batch size attached.
  • x86 runs correctly but without the main optimisation. The build is verified equivalent; clang vectorises it as floats and emits no integer dot product, so x86 gets the pre-SDOT arithmetic. A vpdpbusd path is the obvious fix and is not implemented. No native x86 timing exists for this engine.
  • Evaluated on one machine, an Apple M4 Max. Independent replication welcome, and especially welcome on non-ARM hardware.

What was tried and did not work

Recorded so it is not rebuilt from the same reasoning.

approach result
x86 AVX-VNNI matmul path not attempted: a VNNI binary will not execute under Rosetta, so it could not be verified on the available hardware. Named as headroom rather than guessed at
Patch width 4 (more CTC frames) 92.2% against patch-8's 97.8%; each token sees too little context
Dictionary word repair built three times; detection works (greedy-vs-beam disagreement is an 11.6× error signal), correction does not
Larger LM corpus saturated at 4,933 of 9,025 contexts
Caching ternary weights as float32 1.13× faster end to end, +4 MB RAM
Per-head attention no memory saved; CPython does not return freed memory to the OS
Per-line stretch from frames-per-character the obvious adaptive rule. +0.25 exact-line points against an +11.60 oracle — prediction fails, and fusion gets the same gain by not choosing
Grammar-constrained decoding 41% of edits are deletions, which no grammar can reach, and 86% of the reachable substitutions stay inside their own token class
Turning the LM off on short lines a 59-line pilot said +5.08; 2,000 lines said −2.67. The pilot was noise
White margin around short lines 8px costs 1.3 points, 16px collapses exact-line match to 1.07% — the model reads the margin as a space
String voting (ROVER) instead of fusion 78.05% against fusion's 79.00%, and slower
-O3 / -march=native / -funroll-loops all within noise of -O2 on this CPU
8-way output unroll slower than 4-way (15.5 against 15.0 ms/line): 8 accumulators plus 8 weight pointers spill the register file
Thread-local scratch pointer 34x slower (498 ms/line). The pointer itself lived in TLS, so every element access in the inner loop paid a thread-local resolution. Reading it once into a plain local fixes it -- and then the change is worth 0.2%, so the allocator was never the bottleneck
Self-calibrating confidence gate a 128-line warm-up replaces the hardcoded threshold, but the answer then depends on the ORDER lines arrive in: the same 2,000 lines reversed scored 75.95% against 76.10%. A batch tool whose output changes with input order cannot be diffed between runs
x86 AVX-VNNI matmul path still not attempted: a VNNI binary will not execute under Rosetta, so it could not be verified on the available hardware

Team C · experimental research release · v4 (threaded, register-blocked C engine) · model weights unchanged since v3