Architecture¶
Transcribe keeps a small ownership model on purpose.
Shape¶
Streamlit UI (8510) ──┐
├──► services (project, ingest, job, export, archive, doctor)
CLI ──────────────────┘ │
▼
workspace (data/)
├── corpus/ # corpus-index + import-runs (bulk import)
├── projects/<…>/ # one managed notebook directory each
│ ├── project.json
│ ├── sources/ + pages/ renders
│ ├── results/<page_id>.json
│ ├── analysis/ (optional until first analysis artifact)
│ ├── detection/ (optional until first detection artifact)
│ └── page_metrics/ (optional until first ink/blankness publish)
└── cache/archive.sqlite (disposable)
Ownership boundaries¶
Concern |
Owner |
|---|---|
Corpus identity, notebook order, workspace locks |
|
Managed originals, fingerprints, duplicates |
|
Bulk ImportRun / plan / resume |
|
Corpus/notebook integrity + bulk-import acceptance gate |
|
Durable notebook directory layout + per-notebook journal |
|
OCR generations + edits |
Per-page results — contracts/page-result.md |
Analysis inputs / results / storage / eligibility |
contracts/analysis-document.md · analysis-result.md · analysis-run-storage.md · notebook-eligibility.md |
Prompt definitions / detection findings / detection runs |
contracts/prompt-definition.md · contracts/detection-definition.md · contracts/detection-finding.md · contracts/detection-run-storage.md |
Page ink / blankness / hue metrics |
|
Portable interchange |
Export snapshot — contracts/notebook-export.md |
OCR HTTP |
|
UI widgets |
|
Organisation tag catalog (slugs, labels, colours) |
|
Workspace search/timeline |
|
Key runtime objects (shape, not schema)¶
ProjectService — load/save settings and metadata with load→modify→validate→write under the mutation lock; reconciles interrupted attempts when the job lock is free
IngestService — stages, journals, promotes, then commits the manifest; recovers incomplete journals on open/load
JobCoordinator / JobPlan — freezes model identity, prompt, preprocess, options, targets, provider binding, and optional OCR cleanup identity (mode/model/digest/validator policy) at job start; workers consume the plan, not live UI settings. Multipass reuses frozen single-model plans with
activate=false/pass_id, then rank + composite (contracts/ocr-multipass.md). Per frozen vision plan: after 3 consecutive timeouts or 1 fatal model-load error (unknown model architecture/ loader crash), remaining pages are skipped (circuit_open) so a hung or unloadable model does not burn the notebook; multipass continues with remaining models.OcrBatchRun — durable multi-notebook OCR batch (contracts/ocr-batch-run.md); UI Workflow → Transcribe → Batch
AnalysisBatchRun — durable multi-notebook Analyse batch (contracts/analysis-batch-run.md); UI Workflow → Analyse → Batch (orchestration only; publish stays per-notebook)
ExportService — one coherent snapshot, then multi-format promote
ArchiveService — disposable FTS cache with WAL/busy timeout and delete-and-rebuild on corruption; cheap TTL short-circuit uses a workspace mutation-generation token (callers bump after project mutations)
DoctorService — structural integrity (+ optional deep hashing); quarantined ingest journals reported as errors
CorpusDoctorService / CorpusIndexStore / ImportRunStore — workspace corpus authority under
data/corpus/(runtime-normative; see corpus contracts)AnalysisCoordinator / AnalysisRunPlan / AnalysisRunner / AnalysisStorage — project-scoped async batch runs freeze an
AnalysisRunPlan(modules, EffectiveConfig, text-model identity) and execute under.transcribe.analysis.lock; publish underanalysis/; UI freshness viamodule_freshness/planned_cache_identity(UI must not hand-build cache identities). Mid-run settings apply to the next run only; crash/reopen marks orphaned attempts/runsinterruptedwithout clobbering published resultsDetectionRunner / DetectionStorage / prompt_engine / Prompt Hub — prompt-backed page/window detectors (
poetry,todo_lists,lists,quotations,beer_labels, custom) plus lexical counters (first_person,swear_words) andnames(people from NER); publish findings underdetection/; Settings → Prompts resolves OCR/cleanup/detection definitions with workspace overrides; freshness viadetector_freshness/ planned cache identity. Opt-in auto-tag unionsfinding_type(or detected person names fornames) onto span pages and is not part of detector cache identityTagService / TagCatalogStore — workspace
personal_corpus.tag-catalogatdata/config/tag-catalog.json; assignments remaintags: string[]on notebooks/pages; host-agnostic kernel intranscribe.tagging.kernelfor a future TranscriptX copyPageMetricsService — Pillow ink coverage / blankness / dominant hue over active renders; publish under
page_metrics/; cache identity = algorithm version + ordered(page_id, render_sha256)(not text Analyse)Visual declutter — Pillow scanner-border crop at import and via
ProjectService.reapply_visual_declutter(Settings → Configuration); provenance on renders (contracts/source-asset.md)Ollama discovery cache — thread-safe model metadata keyed by normalized base URL + transport timeout; providers stay lightweight execution clients
Explicit non-goals for the core architecture¶
Making SQLite the system of record
Introducing a task queue or multi-process worker fleet for v1
Coupling to TranscriptX libraries
Making a post-1.0 context corpus (photos, chats, Slices) part of v1 — sequenced on ROADMAP.md After 1.0; that programme must not make SQLite the system of record
Creating
data/context/or context locks before 1.0 — absence of context trees remains a valid workspace; future lock order is corpus → context → notebook (documented for After 1.0; not implemented)
Shipped core analysis + deferred / future / out-of-scope dispositions: ROADMAP.md. Path to 0.9.0 / 0.9-1 / 1.0: ROADMAP.md. Docs authority model: dev/CONTRIBUTING.md.