Known limitations¶
Handwriting quality varies. Some vision models return empty text or fail to load. Remote Ollama sends page images off this machine. This page is the honest list — not a promise that OCR will be perfect.
Product promise: PRODUCT.md.
OCR quality¶
Operational guide: runtime/ocr.md.
Handwriting quality varies widely by model, lighting, and page density
Vision model availability and architectures differ across Ollama builds (a listed “vision” model may still fail to load)
Preprocess default is none;
gentle_contrastis optional and Pillow-based (no OpenCV in v1)Visual declutter (import-time, separate from OCR preprocess) defaults on (
ingest.visual_declutter_enabled). Ships grey/light-grey scanner-bed crop, stark-white overscan/gutter crop, and residual rounded-corner bed wedges; detection is conservative (many pages no-op). Failures fall back to the pre-declutter PNG and never fail import. Changing declutter settings alone does not rewrite existing notebooks — use Settings → Configuration → Re-apply visual declutter, Review → Cleanup (this page or all pages), or a new import to crop existing pages. Re-apply does not re-run OCR and cannot restore already-cropped margins when turned off.Page ink / blankness metrics (Review strip + Overview rollup) are approximate Pillow heuristics over the active render. The notebook’s explicit
cover_page_idis omitted (not measured or shown). Ruled lines, shadows, stains, and colour casts can inflate “ink”; hue labels (black/blue/ …) are coarse peaks, not calibrated colour science. Metrics invalidate when active render bytes change; they are not Analyse text modules and do not affect OCR.Optional OCR cleanup (Run tab /
--cleanup) adds a second text-model Ollama call per page after vision OCR. This can materially increase latency, memory use, and Ollama contention. Cleanup runs sequentially on the page worker after OCR (no extra parallelism). Failures and validator rejections keep raw OCR and never fail the page; rejected model output is discardedCleanup sends OCR text (not page images) to the configured Ollama host; remote hosts still exfiltrate that text by design of that configuration
Vision OCR always sends a
num_predictcap (default 4096). That stops a looping generate from running until the HTTP timeout. Hitting the cap recordstruncatedin allowlisted provider metadata and does not fail the page. Defaultnum_predictis omitted from skip fingerprints so existing attempts still matchEmpty OCR (whitespace-only model output) is failed (
empty_output) and does not replace a prior succeeded reading. Review Failed OCR is the queue; historical emptysucceededattempts are repaired on ReviewThinking vision models (for example
gemma4,qwen3-vl,gpt-oss) often consume the fullnum_predictbudget in hidden reasoning and return no text — the job showsempty_outputon most pages. Transcribe excludes these from vision pickers; prefer OCR-oriented or probed tags — ocr_model_matrix.mdDeepSeek-OCR (and similar recipe tags) ignore long faithful instructions and often emit one token then stop. Transcribe applies a short
free_ocrprompt unless you set a custom prompt — ocr_model_recipes.mdOllama generate timeouts are not retried. Connection errors and 5xx responses still retry (3 attempts). A hang therefore fails in one HTTP timeout (default 300s), not ~15 minutes
General vision-language models (for example
llava) can hang or time out on dense notebook scans even when listed as vision-capable. Prefer OCR-oriented tags for handwriting. After 3 consecutive timeouts on one frozen vision plan, remaining pages for that model are skipped (progresscircuit_open); the job record stayscompletedand the UI shows Completed with gaps. Single-modeltranscribe runexits 1 so automation does not treat the notebook as fully transcribed. A multipass compare continues with the next modelSome Ollama “vision” tags still fail to load on a given build (example:
llama3.2-vision:11b→unknown model architecture: 'mllama'on Ollama 0.32.x — see ollama#16547). Transcribe classifies these as non-retriablemodel_loaderrors and skips remaining pages for that model after the first failure (samecircuit_openpath as timeouts). Prefer a working alternate family (for examplegranite3.2-visionorminicpm-v) until the host Ollama/model pair loads cleanly. Batch OCR inherits the same per-notebook circuit (one bad model does not keep calling every page in that notebook).Multipass compare runs each selected vision model across the notebook, then a text-model rank (text-only v1) and optional composite merge. Cost scales with model count × pages plus rank/composite calls; on Batch multipass it also scales with notebook count. Composite is assistive, not ground truth. Rank failure falls back to chronological attempt order in Review. Vision phases default cleanup off (CLI
--cleanup/ UI “Clean OCR during compare” to opt in). The UI starts compare in a background thread like single-model Start; Stop cancels remaining pages of the current model and remaining models (and does not start remaining batch notebooks). Rank/composite still run for pages that already have ≥2 succeeded vision attempts
Import / PDF¶
Encrypted PDFs are rejected
Very large sources/PDFs fail closed on configured byte/page/render budgets
PDF rendering uses PyMuPDF; unusual PDF constructs may render poorly
After corpus index recovery, retained quarantine artifacts under
data/corpus/quarantine/are doctor warnings (corpus_quarantine_present) until an operator deletes them — they do not block a healthy corpus
Jobs and identity¶
Fingerprint skip requires verified model identity (digest from Ollama discovery). Unverified tags are always re-run
Model pickers filter discovery: vision/OCR selectors show OCR-appropriate VLMs only; text selectors show completion LLMs only (see ocr_model_matrix.md)
Cancelling stops scheduling after the current page; in-flight pages still finish. The progress panel shows Cancelled (not Failed). During compare, remaining vision models are not started; rank/composite still run for pages that already have ≥2 succeeded vision attempts
Mid-job settings changes apply to the next job only
Archive / cache¶
Workspace search/timeline depends on a rebuildable SQLite cache. Corrupt or incompatible caches are deleted and rebuilt
Cheap
ensure_indexshort-circuit uses an explicit mutation generation token (data/cache/archive.generation), bumped after import/OCR/edit/metadata — not directory mtimes (in-place result edits do not reliably change dir mtime). Per-project rebuild signatures still use result file mtimes inside a rebuildAuto-suggested / inherited page dates (unapproved) still index in the Library timeline; approval status is not a filter. Review states this in-product and offers batch approve/ignore for suggestions.
Diary date auto-extract (early page text) understands compact
YYMMDD,DD/MM/YYYY,DD/MM/YY,YYYY-MM-DD, and English month names (Jan 2, 2018). Ambiguous numerics are day/month (DMY). Time-of-day is ignored. OCR can still garble stamps; pages that look stamped but fail to parse stay undated (no inheritance) until ReviewOllama model discovery metadata is cached by base URL + transport timeout; Refresh invalidates. Execution clients stay lightweight. Model information shows verified vs unverified identity (digest) and preference last-used when available.
Library Covers grid paging defaults to show all (
ui.archive_notebooks_initial = 0). A positive value loads that many cards before Show more; session state can expand further until rerun/reset.Reading/Review Thumbnails grids serve small disposable JPEGs (
.cache/thumbs/*.grid.jpg, max edge 128). Import (single-notebook and bulk) warms cover + grid thumbs; Settings → Configuration → Thumbnails → Regenerate thumbnails force-rewrites them for an existing notebook. Older notebooks without grid thumbs generate on first Thumbnails open (one-time cost per page).
Workspace backup / restore¶
Operator guide: backup_and_restore.md.
Full-workspace ZIPs pack notebooks, corpus, and config (optional inbox/exports). They never treat
data/cache/archive.sqliteas authority; restore deletesdata/cache/so Archive rebuilds.Restore is replace-only onto the current
TRANSCRIBE_*mounts (path-agnostic role-root layout). There is no merge / per-notebook restore in v1.Large workspaces: use CLI disk paths (
backup create/restore); the Settings UI does not download or upload multi-GB ZIPs through the browser.Backup ZIPs contain page images and OCR/analysis text — treat them as sensitive local files (no encryption; no cloud upload from Transcribe).
Create/restore refuse while a corpus lock or notebook OCR/analysis job lock is held; create refuses clobbering an existing dest without
--force; restore refuses archives that sit under trees being replaced (except{EXPORT}/backups/).Automatic pre-restore safety ZIPs use default create options (inbox/exports omitted). If restore fails mid-replace, recover from that
pre-restore-*.zip(path is included in the error when written).Operator guide: backup_and_restore.md.
Privacy¶
Local-by-default Ollama. Remote hosts exfiltrate page images by design of that configuration
Transcribe does not ship cloud OCR providers
Analysis¶
Operational guide: runtime/analysis.md. Settings / presets: runtime/settings.md.
Core analysis modules are shipped; quality follows OCR text quality (noisy handwriting hurts NER, topics, and LLM grounding)
Prefer OCR cleanup / second-pass LLM verification and human review edits to improve text before analysis; a dedicated
ocr_qualityanalysis module is deferred (ROADMAP.md)Optional extras (
bertopic, spaCy NER path, fine-grained emotion) degrade to named capabilities (unavailable_extra) rather than silent substitutes. The names detector depends on that NER path.LLM Summaries / Ask notebook need a text Ollama model (workspace default, batch pick, or per-notebook); missing model →
unavailable_model. Deterministichighlights→summary→insightsstill work offlineBatch Analyse runs from the preset form (plus a shared text-model pick when LLM modules are included); View pages are read-models over
published.json. Ask notebook remains an ad-hoc actionAnalyse → Batch runs the same frozen plan template sequentially across notebooks (dual progress bars: notebooks + modules). Empty-text notebooks are skipped; there is no OCR-style Force flag. This is orchestration only — not cross-notebook / corpus-level Analyse
Analyse → Batch Pick notebooks labels show published presence (
no analysis,existing analysis,existing degraded analysis, failed / interrupted / running). They do not scan content freshness; use Notebooks needing analysis for out-of-date resultsBatch runs use a frozen
AnalysisRunPlanunder a project analysis lock; mid-run settings / text-model / module-list changes apply to the next run onlyStreamlit UI interruption does not drop an in-process batch (AnalysisCoordinator). Process crash/reopen marks orphaned attempts and run records
interruptedwithout clobbering published results; re-run uses cache hits — no auto-resumeFreshness is computed via
module_freshness/ planned cache identity — not hand-built identities in the UIView consume pages share derived
AnalysisHealth(samecontent_revision+ aggregate rules); Ask notebook remains ad-hoc and does not update batch healthBatch launches freeze an
AnalysisRunPlanwithplan_hashat confirm; start refuses hash mismatch and does not re-snapshot settingsNamed presets carry
content_version(bumped on Settings save); runs record preset identityExports stamp notebook
content_revisionon JSON, manifest, Markdown, and plain textDedicated Patterns tab is not shipped; payloads feed Themes instead (optional polish under the usability wave, not deferred reinterpretation modules — usability_wave_plan.md)
People & Places (View): People and Places sections each toggle This notebook | All notebooks. Geocoding via OpenStreetMap Nominatim is opt-in and cached under
data/cache/geocode.jsonOverview / Mood corpus or period compare averages other notebooks’ published numeric metrics (this notebook excluded). Year / date-range use diary
date_start/date_end; undated notebooks count only under “Entire corpus”. Peers without a published result for that module are skipped — charts need at least one peer with dataWord themes offer Basic (static frequency cloud) or Advanced (interactive explorer with search / top N / min value / sort / CSV — TranscriptX explorer controls). Advanced uses a vendored
wordcloud2.js(offline). Basic uses the defaultwordcloudpackage. Analysis still stores frequencies only — images/explorer state are not durable artifacts.Deferred reinterpretation modules are not scheduled; product focus is the usability wave (trust, Analyse product UX, first-run, daily workbench) for the shipped surfaces — ROADMAP.md Now
Analysis results live under project-local
analysis/and invalidate with text/config/parent changes — see contracts under CONTRACT_INDEX.md
Integration¶
No TranscriptX dependency. Future notebook handoff is documented separately and is not shipped behaviour: INTEGRATION_SEAM.md