Local LLM integration (Ollama)¶
Optional local AI stays on your machine unless you point it at a remote URL. Transcript text can appear in prompts — enable it only when that is acceptable. Modules that need a live model are off by default.
LLM-backed analysis modules are disabled by default. Enable them only when you have a local Ollama daemon (or another configured endpoint) and understand that prompts may contain sensitive transcript content.
Configuration¶
Environment variables (highest precedence):
export TRANSCRIPTX_LLM_ENABLED=1
export TRANSCRIPTX_LLM_PROVIDER=ollama
export TRANSCRIPTX_LLM_MODEL=qwen3:8b
export TRANSCRIPTX_LLM_BASE_URL=http://localhost:11434
export TRANSCRIPTX_LLM_SEED=42
Or in config.json:
{
"llm": {
"enabled": true,
"provider": "ollama",
"model": "qwen3:8b",
"base_url": "http://localhost:11434",
"seed": 42,
"request_timeout": 1350,
"max_input_chars": 48000,
"default_temperature": 0.3
}
}
Configure timeout via llm.request_timeout / TRANSCRIPTX_LLM_REQUEST_TIMEOUT (default 1350 seconds / 22.5 minutes).
Remote endpoints: the default base_url is local. A non-local URL sends transcript content to that endpoint. TranscriptX does not block remote URLs but you are responsible for where your data is sent.
Run Analysis model selection¶
On Run Analysis (Transcript, Group, and Batch), the LLM setup section appears only when the effective module list includes a live-LLM consumer (registry requires_llm modules that still need a live call), or when group analysis has group_llm_synthesis enabled. When llm.enabled and provider=ollama, that section shows the project-active Model preset and optional Change for this run overrides (shared model or per-module), including a Model information expander with installed-tag guidance. Non-LLM presets hide the section entirely so model overrides are not offered for irrelevant runs.
Settings → Models owns management actions:
Refresh installed Ollama tags
Model information / guidance table
Create / overwrite Model presets (ProfileManager target
llm_models, under.transcriptx/profiles/llm_models/)Set the project-active Model preset
Per-run selections are snapshotted onto the analysis request and do not rewrite llm.model unless you save a Model preset under Settings → Models (or activate one under Settings → Configuration → Active Profiles).
Resolution precedence for each LLM consumer (narrative_summary, llm_summary, llm_speaker_summary, llm_action_items, llm_custom_qa, chart_descriptions, group_llm_synthesis):
Request override from Run Analysis / Batch
Active
llm_modelsprofile applied ontollm.model_selectionGlobal
llm.model/ default (qwen3:8b)
Effort-profile model fields are not part of this chain when a consumer id is set. Corrections Studio (no consumer id) may still use an effort-profile model over global llm.model.
On the run form, Project default reflects the already-applied project llm.model_selection pack. Custom (this run) (under Change for this run) keeps free-edited widgets for this launch only. Unavailable saved tags are cleared with an explanation (no silent substitute); launch stays gated until an installed model is chosen when LLM modules are selected.
If LLM is disabled or the provider is not Ollama while selected modules (or enabled group synthesis) need LLM, the launch button stays disabled. Non-LLM analysis remains runnable when no live-LLM modules are in the effective module list.
Thinking models (JSON-unsafe): tags matching qwen3* (including qwen3.8 / qwen3.6), deepseek-r1*, and gpt-oss* often put tokens in Ollama’s thinking field and leave response empty when TranscriptX requests format=json. That fails narrative_summary, llm_action_items, chart_descriptions, and group_llm_synthesis. The installed list is live from Ollama (/api/tags, short cache; Refresh models under Settings → Models). Settings → Models shows every installed tag in the shared picker. The compact Run Analysis selector hides thinking tags from shared picks whenever any JSON module is selected, and from per-module rows for JSON consumers, and names the omitted tags under the picker. Launch stays gated if a saved preset still assigns a thinking tag to a JSON consumer. Prefer non-thinking tags such as gemma3:*, qwen2.5:*, llama3.2:*, mistral:*, or mistral-nemo for those modules (plain-text llm_summary / llm_speaker_summary may still work with thinking models).
Complementary Ollama picks for transcript analysis¶
Useful additions when the local library is already Qwen-heavy (or JSON modules need non-thinking tags):
Tag |
Size class |
Why pull it |
Best TranscriptX fit |
|---|---|---|---|
|
small |
Fast non-thinking captions with 128K context; better than toy 1B tags |
|
|
mid |
12B instruct with extreme context headroom; stable JSON |
Long-session |
|
large |
Multilingual / high-quality non-thinking alternative to large Qwen/GPT-OSS |
Shared high-quality JSON + prose ( |
ollama pull gemma3:4b
ollama pull mistral-nemo
ollama pull gemma3:27b
Under Change for this run, the Model information expander describes installed tags (including these) for per-module assignment.
If a selected model is missing at generate time, the LLM consumer fails with a clear model-missing error (no silent substitute). Corrections Studio continues to use global llm.model only.
Modules¶
Module |
Description |
Depends on |
|---|---|---|
|
Grounded executive narrative from deterministic |
|
|
Abstractive transcript summary from readable transcript text (plain text, |
segments only |
|
Abstractive summary per named speaker from that speaker’s utterances only (plain text, |
segments + speaker labels |
|
Structured meeting extracts (decision / commitment / action_item / proposal / open_question) with owner, deadline, status, quote (strict JSON, |
segments + speaker labels |
|
Answers to user-defined questions with grounded citations (strict JSON envelope+rows; empty questions succeed without Ollama) |
segments |
|
Finalize-phase per-chart LLM narratives (temperature 0.0, JSON). Excluded from DAG execution; run by the run-finalization coordinator after all charts exist |
selected + |
All LLM modules except finalize-phase chart_descriptions are included in the recommended default module list when enabled. chart_descriptions is selectable and recommended with other LLM modules but executes only in finalization. Uncheck Use recommended modules in Run Analysis (or pass an explicit modules list via the API) to opt out. When LLM is disabled they are skipped before execution (not failed) with reason LLM disabled. For chart_descriptions, a skipped generation is still committed (ACTIVE + epoch) with zero client calls.
analysis.chart_descriptions¶
{
"analysis": {
"chart_descriptions": {
"enabled": true,
"chart_set": "all"
}
}
}
enabled(defaulttrue): gate with module selection and LLM enabled.chart_set:all(default) |transcript_group|overview_only.Layout:
.chart_descriptions/LATEST_ATTEMPT.json,ACTIVE.json,generations/<id>/with COMMIT/index/outcome/descriptions.Resolvers require ACTIVE.attempt_epoch to match LATEST_ATTEMPT (suppresses stale text after crash).
Ordering: chart-description publish → group LLM synthesis → single manifest write under one run-finalization lock.
Privacy: transcript text for LLM modules (llm_summary, llm_speaker_summary, llm_action_items, and structured input to narrative_summary) is sent only through the configured Ollama path (llm.provider=ollama). TranscriptX does not block non-local base_url values, but you are responsible for where your data is sent; the default is loopback (http://localhost:11434). Group LLM synthesis sends member summary texts only (not raw transcripts) over the same Ollama path.
llm_speaker_summary is skipped when the transcript has no eligible named speakers (for example, before Speaker ID mapping). Ignored speakers are excluded.
Selecting narrative_summary automatically runs the summary dependency chain first.
Group LLM synthesis¶
After group finalize collects per-member llm_summary / llm_speaker_summary artifacts, an optional cross-session synthesizer writes generation-scoped rollups under .group_llm_synthesis/ (ACTIVE/COMMIT). See group_llm_synthesis_contract.md.
{
"analysis": {
"group_llm_synthesis": {
"enabled": true,
"effort": "high"
}
}
}
enabled(defaulttrue): when false, finalize commits a skipped generation and flips ACTIVE so prior success is no longer shown.effort: same tiers asllm_summary. Resolves effectivemax_input_chars/request_timeout/max_output_tokensfor synthesis calls only; it does not mutate process-globalllm.*values used by other modules.Ollama-only: non-Ollama or disabled LLM → skipped generation (not a hard group failure).
UI: group Overview/Insights use the central resolver; no member-summary primary fallback on group runs.
llm_summary effort (not llm.effort)¶
The setting analysis.llm_summary.effort controls summary effort for the llm_summary module only (full-transcript abstractive summary). It is not llm.effort and does not affect narrative_summary, global llm.* provider settings, or other analysis modules.
Valid values: low, medium, high, max (default: high).
{
"analysis": {
"llm_summary": {
"effort": "medium"
}
}
}
When llm.enabled is true and llm.provider is ollama, llm_summary resolves builtin effort profiles that set max_input_chars, request_timeout, and max_output_tokens for that module run. Those effort limits replace the corresponding llm.* values for llm_summary only; they are not merged with user-tuned llm.max_input_chars / llm.request_timeout / llm.max_output_tokens. The model still defaults to llm.model unless a future profile sets an override.
Tier intent:
low— useful preview modemedium— default completeness-oriented mode for normal long transcripts (not legacyllm.*defaults)high— patient mode for long meetings, workshops, lectures, and dense transcriptsmax— push-the-laptop mode; prefers waiting over truncating or timing out
This pass supports provider="ollama" only for llm_summary, llm_speaker_summary, and llm_action_items; other non-null providers raise configuration errors. Shared eligibility is enforced via require_ollama_analysis in core/analysis/llm_support/runtime.py, which also owns the shared effort-profile map (resolve_llm_runtime), the analysis Ollama client factory (build_ollama_analysis_client), and input-coverage provenance (build_input_coverage). These helpers are shared by all three transcript-direct modules; there are no summary-specific aliases.
Artifact provenance records effort, effort_profile, resolved limits, and input coverage (input_truncated, input_chars_total, input_chars_used, input_coverage_ratio). The legacy input_chars field remains the full prompt size including wrapper text.
llm_speaker_summary effort¶
The setting analysis.llm_speaker_summary.effort controls per-speaker summary effort for the llm_speaker_summary module only. It uses the same builtin tiers as llm_summary (low, medium, high, max; default high in code).
{
"analysis": {
"llm_speaker_summary": {
"effort": "high"
}
}
}
Each named speaker triggers one sequential Ollama call over that speaker’s utterances. Per-speaker failures (for example, an empty model response) are recorded in the index artifact; the module succeeds when at least one speaker summary is written.
Artifacts:
llm_speaker_summary/data/speakers/{base}_{Speaker}_llm_speaker_summary.json(+.md) per speakerllm_speaker_summary/data/global/{base}_llm_speaker_summary_index.json(+.md) listing speakers and statuses
llm_custom_qa (custom questions)¶
Answers user-defined questions against a bounded transcript excerpt (and, when structured execution is on, routed evidence packs). Product entry: Settings → Questions and Run Analysis / Batch question picker. Insights: global answers under the summary hero; per-speaker answers under speaker summaries when present.
Live path vs structured path¶
There is one public schema: transcriptx.llm_custom_qa.v1. There are
two implementation paths in code — not two product schemas:
Live (default) |
Structured (off) |
|
|---|---|---|
When it runs |
Always, unless structured is explicitly enabled |
Only if |
Code |
|
|
Questions |
Text list projected from the library/picker |
Full structured questions ( |
Commit layout |
Alias files ( |
Generation-named files ( |
Contract module |
|
|
Why structured is off: the richer path (per-question scopes, evidence-pack
routing, generation-named commits) is implemented and covered by tests, but it
is not flipped on for production runs yet. Shipping stays on the live path so
behavior stays stable. Readers already accept structured-shaped artifacts when
they appear (e.g. question_order present); writers do not produce them until
the gate is turned on.
Do not invent a second schema_id for structured — both paths stamp
SCHEMA_ID / MODULE_VERSION from versioning.py.
Settings¶
{
"analysis": {
"llm_custom_qa": {
"effort": "high",
"saved_questions": [
{"text": "What was decided?", "scopes": {"global": true, "per_speaker": false}}
],
"evidence_pack_ids": null,
"include_transcript": true,
"routing_enabled": true,
"max_library_questions": 50,
"max_library_total_question_chars": 20000,
"max_questions_per_run": 8,
"max_question_chars": 500,
"max_run_total_question_chars": 4000,
"max_answer_chars": 800
}
}
}
saved_questions is structured (text + scopes); plain list[str] is
accepted and treated as global-only. Persisted in project config.json under
CONFIG_DIR (Docker: /data/.transcriptx, host mount via HOST_CONFIG_DIR).
Prefer HOST_CONFIG_DIR outside the git clone so the library survives wiping
./data. The web app hydrates get_config() from that file on startup.
evidence_pack_ids: null means all present and
future catalog packs (expanded only onto the run plan). Request field
llm_custom_qa_questions: omit/null → library; [] → empty-run success;
list (str or structured) → request. Resolve/bind only when the module is selected.
The Run Analysis / Batch picker treats an empty selection as implicit skip
(omit); the empty-run ([]) path stays API-only for power users.
Contract¶
Artifact
schema_idistranscriptx.llm_custom_qa.v1Corpus capped (
MAX_CUSTOM_QA_CORPUS_CHARS), tail-preferring when truncatedLive commits use bare aliases; structured commits (when enabled) use generation files
{stem}.json.{gid}/.md.{gid}, commit markers (run_execution_id), atomic.active, and best-effort aliasesReaders resolve canonical stem then active→marker; structured payloads with
question_orderare validated viastructured_contracts.pyEmpty questions /
no_scheduled_cellsstill succeed and commit
Artifacts¶
Live / alias:
{base}_llm_custom_qa.json(+.md)Structured / authoritative (when enabled):
…/{base}_llm_custom_qa.json.{generation_id}(+.md.{gid})Generation metadata:
{base}_llm_custom_qa.questions_metadata.{gid}.jsonGroup:
qa_answer_rows.json,qa_member_failures.json(via group content loader)
llm_action_items effort¶
The setting analysis.llm_action_items.effort controls meeting-extract effort for the llm_action_items module only. It uses the same builtin tiers as llm_summary (low, medium, high, max; default max because meeting-extract JSON lists often truncate under lower output budgets).
{
"analysis": {
"llm_action_items": {
"effort": "max",
"coerce_v1_artifacts": false
}
}
}
coerce_v1_artifacts (default false): when true, in-memory coerce legacy v1 artifacts to record_type=action_item with provenance.compat=v1_coerced for presentation/group aggregation. Without it, mixed-version group members fail explicitly and UI shows a legacy v1 path — never as native v2. Does not rewrite on-disk artifacts.
Artifacts:
llm_action_items/data/global/{base}_llm_action_items.json(+.md)
Output contract (v2)¶
schema_id: transcriptx.llm_action_items.v2
render_contract_id: transcriptx.llm_action_items.render.v2
module_version: 2 · prompt_version: 7
Field |
Notes |
|---|---|
|
Ordered meeting extracts; empty list is a successful result |
|
|
|
Non-empty trimmed description |
|
Verbatim transcript wording or |
|
|
|
Exact transcript substring after whitespace normalisation, or |
|
Finite float in |
|
Includes |
|
Includes |
When output_truncated is set, Insights / Overview show a warning that extracts may be incomplete and that re-running unchanged will usually repeat the failure — raise analysis.llm_action_items.effort to max and/or pick a stronger JSON-capable model for llm_action_items, then re-run only that module. Hard parse failures persist the same remediation text into run_results / block availability (instead of a bare “Run the required analysis modules”).
When the artifact has an empty items list but diagnostics.items_raw > 0 (typically items_invalid_dropped / status / grounding drops), Insights / Overview show a warning that records were returned then discarded, and steer users toward mid+ JSON-capable models. A genuine empty model response (items_raw == 0) still shows the short “No meeting extracts found” caption. Empty-extract runs also write {base}_llm_action_items.raw.txt beside the JSON for parse debugging.
Top-level malformed JSON / missing items / oversize raw output fail with llm_invalid_response. Unknown wrapper keys around items are ignored. Per-record isolation: invalid items are dropped with diagnostics so sibling valid extracts survive; unknown per-item keys are stripped rather than dropping the record. Grounding is record-type agnostic and never rewrites record_type / text / owner / deadline / status. Bounds: max 48 total / 16 per type after validation. Markdown uses section labels from RECORD_TYPE_LABELS and always includes AI-generated draft. Human review required.
Group aggregation schema_version: 2 with fixed count_<record_type> and status_* session columns.
Identity for caching is a distinct namespace (provenance.cache_key); schema/prompt/module bumps invalidate the key.
UI and export¶
Insights layout (
default,executive): blockllm_action_items_blockrenders sectioned Markdown or a typed table titled Meeting extracts.Overview module metrics: summary extractor surfaces meeting-extract counts by type and status.
Zip export: JSON/MD are included in module/data exports; Overview ZIP
index.html/index.epublist a Meeting extracts summary section when those artifacts are selected (seeresolve_export_text_summariesintranscriptx.export.resolve_summaries). Full export behaviour: runtime/export.md.
Truncation¶
llm_summary and other transcript-direct Ollama modules (including llm_action_items) cap the full user prompt (instructions, delimiters, and transcript block) to an input budget using the existing head/tail truncation algorithm. On the Ollama effort path, the budget comes from the selected effort profile’s max_input_chars. When the formatted transcript exceeds the budget, TranscriptX uses a deterministic head/tail strategy:
Reserve space for an omission marker in the middle.
Allocate roughly 60% of the remaining budget to early segments and 40% to late segments (whole segments only).
Long transcripts may omit middle material; provenance records
truncated, segment counts, andtruncation_strategy.
Error codes¶
Failed LLM modules return execution_status=failed with a stable error_code in module_result and run_results.json (module_outcomes):
Code |
Meaning |
|---|---|
|
Daemon unreachable after retries |
|
Configured model not installed |
|
Request timed out after retries |
|
Malformed JSON, wrong shape, empty response, or invalid narrative JSON |
|
Non-retryable HTTP/client generation failure (e.g. HTTP 400/401/403, exhausted 5xx retries) |
|
Invalid LLM configuration |
|
Required upstream module output missing, failed, skipped, or blocked |
|
No usable transcript or summary signal |
llm_invalid_response is reserved for successful HTTP responses with unusable body content. HTTP 4xx/5xx generation failures map to llm_generation_error unless the response body explicitly indicates a missing model (llm_model_missing). A bare HTTP 404 with an empty body maps to llm_generation_error, not llm_model_missing.
llm_dependency_missing may include structured error_context on the module error envelope, for example {"dependency": "summary", "state": "missing|skipped|blocked|failed"}. UI and canonical outcome rows surface error_code alongside human-readable messages.
Prompt-budget validation happens at two levels:
Config load (global):
llm.max_input_charsmust be at least the fixed prompt-envelope minimum (delimiters and safety copy only, no feature instruction;core/llm/prompting.py::prompt_envelope_min_chars). Config load rejects lower values.Runtime (per feature): because effort-profile limits replace the global
llm.max_input_chars, each transcript-direct module (llm_summary,llm_speaker_summary,llm_action_items) validates its resolved effort budget against its exact instruction plus delimiters (core/llm/prompting.py::require_prompt_budget) before constructing a client or making a network call.narrative_summaryis excluded: it uses a findings-rewrite prompt, not the bounded transcript envelope.
Provenance¶
Successful LLM artifacts include mandatory provenance fields:
llm_request_sha256— SHA-256 of canonical JSON{user, system?}sent toclient.generate()model,provider,seed,temperature,max_output_tokens,generation_options(including effectivenum_predict)transcriptx_versionwhen importableFor
llm_action_items: alsomodule_version,prompt_version, andcache_key(distinct cache identity namespace)
Optional metadata such as model_digest is included only when already cached (e.g. from a prior is_available() tags fetch); no extra /api/tags call is made solely for provenance.
Artifact writes¶
LLM modules write JSON/Markdown pairs through core/analysis/llm_support/artifacts.py under an atomic pair promotion with rollback, then registration contract:
Both files are fully staged (and prior canonical files backed up) in a per-write
.staging/subdirectory before any promotion.JSON is promoted first, then Markdown. If the Markdown promotion fails, the JSON promotion is undone exactly once — restored from backup, or removed when there was no prior file. Prior canonical files are never deleted optimistically.
If the rollback itself fails, the original promotion error is propagated (the rollback failure is logged and attached as exception context).
The staging directory is cleaned up in
finally.Artifact registration (
record_file) begins only after both promotions succeed, and there is no filesystem rollback after registration begins. Registration is not transactional: if the first (JSON) registration fails, both files remain promoted and nothing is registered; if the second (Markdown) registration fails, both files remain promoted and the JSON registration remains. Undoing a registration would require anOutputServiceunregister API, which does not exist.
Speaker artifact filenames sanitise display names by replacing spaces and slashes with underscores (llm_support/filenames.py). Distinct names can collide to the same filename (e.g. A B, A_B, and A/B all map to A_B); this is documented, tested behaviour — collision-safe identity is tracked as separate work because changing it would change artifact paths.
Manual smoke test¶
Prerequisites:
Ollama running locally (
ollama serve)Model installed:
ollama pull qwen3:8b
export TRANSCRIPTX_LLM_ENABLED=1
export TRANSCRIPTX_LLM_PROVIDER=ollama
export TRANSCRIPTX_LLM_MODEL=qwen3:8b
python -c "
import os
from pathlib import Path
from transcriptx.app.models.requests import AnalysisRequest
from transcriptx.app.workflows.analysis import run_analysis
os.environ['TRANSCRIPTX_ALLOW_UNMANAGED_TRANSCRIPTS'] = '1'
result = run_analysis(AnalysisRequest(
transcript_path=Path('tests/fixtures/mini_transcript.json'),
modules=['summary', 'narrative_summary', 'llm_summary', 'llm_speaker_summary', 'llm_action_items'],
))
print('success:', result.success)
print('errors:', result.errors)
"
Expected artifacts per module under the run output directory:
narrative_summary/data/global/*_narrative_summary.json(+.md)llm_summary/data/global/*_llm_summary.json(+.md)llm_speaker_summary/data/speakers/*_llm_speaker_summary.json(+.md) per named speakerllm_speaker_summary/data/global/*_llm_speaker_summary_index.json(+.md)llm_action_items/data/global/*_llm_action_items.json(+.md)
Partial failure: if a selected LLM module fails, the overall run is partially failed. Deterministic modules (e.g. summary) and their artifacts remain available. Failed LLM modules produce no canonical LLM artifacts.
Graceful failure when Ollama is stopped: modules report failed with error_code=llm_unavailable after up to 3 connection retries (~2s backoff cap).
Optional live integration tests¶
Client smoke + LLM analysis modules (llm_summary, llm_speaker_summary,
llm_action_items):
export TRANSCRIPTX_LLM_LIVE_TEST=1
# optional: export TRANSCRIPTX_LLM_MODEL=qwen3:8b
# optional: export TRANSCRIPTX_LLM_SMOKE_MODEL=llama3.2:3b
# optional: export TRANSCRIPTX_LLM_ACTION_ITEMS_MODEL=llama3.2:3b
# optional (host-side): export TRANSCRIPTX_LLM_LIVE_BASE_URL=http://127.0.0.1:11434
pytest tests/core/llm/test_ollama_live.py tests/core/analysis/test_llm_modules_live.py -m "integration and requires_api"
Requires a running Ollama daemon and the configured model installed locally.
These tests are excluded from the default fast suite (integration /
requires_api / slow). Module live tests use analysis.*.effort=low to keep
runtime bounded.
On the Mac host, prefer TRANSCRIPTX_LLM_LIVE_BASE_URL=http://127.0.0.1:11434.
A project .env value of http://host.docker.internal:11434 is for the Docker
GUI container and is ignored by these host-side live tests unless you set the
live-specific override.