Run outcome contract

This document defines the canonical run outcome contract for TranscriptX. It is the single source of truth for how run_results.json represents execution outcomes.
Structure and placement of output artifacts are defined in docs/contracts/output-contract-v1.md; this document only defines execution truth and status semantics.

1. Source of truth

  • run_results.json at the run root is the single source of truth for per-module execution outcomes.

  • File presence or absence alone is not truth. Consumers must not infer success/failure/skipped solely from which files exist; they must consult run_results.json.

2. Allowed statuses

The only allowed canonical statuses are:

  • requested

  • enabled

  • blocked

  • skipped

  • failed

  • succeeded

2.1. Status meanings

  • requested: The module was part of the requested plan but not yet enabled (e.g. filtered out before execution).

  • enabled: The module was admitted to the execution plan but did not produce a terminal outcome row in any of modules_run, modules_failed, or modules_skipped.

  • blocked: The module did not run because a dependency or gate blocked it (e.g. missing capability, failed prerequisite).

  • skipped: The module was intentionally skipped despite being otherwise runnable (e.g. user choice, preset behavior).

  • failed: The module attempted to run and reached a terminal error state; downstream modules may have been skipped.

  • succeeded: The module completed successfully according to the pipeline, regardless of whether outputs were cached or freshly computed.

These labels are API-level truth for read-side consumers. Do not rename or add new statuses without updating this contract.

2.2. Run-level run_status (optional)

run_results.json may include a run-level field run_status, distinct from module statuses in §2:

  • running — workspace exists; DAG has not yet written a terminal summary. Module lists under this status are non-terminal (modules_failed must be empty).

  • succeeded | partial | failed | aborted — terminal overwrite after persist (same vocab as orchestrator RunStatus).

Missing run_status on older files means terminal: consumers infer from module lists via the projection in §4. Do not treat file presence alone as success. Stay on schema_version == 1.

The pipeline writes run_status=running after the execution plan and before DAG execute, then overwrites with a terminal value at persist. Diagnostics distinguish live vs interrupted running files via the analysis lock (flock), not mtime.

3. Write-side vs read-side model

  • Write-side model:

    • The pipeline/DAG writes a non-terminal run_results.json (run_status=running) after planning and before execute, then normalizes terminal outcomes and overwrites run_results.json before writing manifests or secondary reports.

    • Persistence order:

      1. Canonical outcome record (run_results.json)

      2. Artifact manifest (manifest.json)

      3. Secondary summaries and projections (reports, derived views)

  • Read-side model:

    • Consumers (GUI, reports, tools) must load typed run outcome context (e.g. via load_run_results / load_run_outcome_context) and derive a status for each module through a single canonical projection path, not ad-hoc logic.

4. Projection rules

Read-side status projection is implemented in run_outcome_truth.py and must follow these high-level rules:

  • modules_run entries project to succeeded.

  • modules_failed entries project to failed.

  • modules_skipped[] entries with execution_status == "blocked" project to blocked.

  • Other modules_skipped[] entries project to skipped.

  • modules_enabled entries not otherwise classified project to enabled.

  • Entries present in run lists but not in the enabled set project to requested.

Cache hits:

  • Cache usage is represented as metadata (used_cache: true, optional reason code such as cache_hit).

  • The status remains succeeded whether or not a cache was used.

5. Schema and loader contract

  • Schema version:

    • Canonical execution truth uses schema_version == 1 (public schema epoch-1).

    • Typed loaders must reject run_results.json with any other schema_version via assert_run_results_schema_supported.

    • Older on-disk summaries must be regenerated by re-running analysis.

  • Typed load path:

    • Consumers must use:

      • load_run_results(...)

      • or load_run_outcome_context(...)

    • and must not parse run_results.json with raw json.load for truth-path consumption.

  • After the schema gate, loaders may apply defaults for missing optional fields (e.g. empty lists, placeholder ids). This does not change the allowed statuses or semantics above.

6. Relationship to manifest.json and other outputs

  • manifest.json is the artifact registry:

    • It records what artifacts were written, where they live, and metadata for reproducibility.

    • It is not the primary source of truth for module status.

  • run_results.json is authoritative for:

    • Which modules ran.

    • Whether each module succeeded, failed, was skipped, or blocked.

    • How cache and gates affected execution.

  • Reports and other derived outputs:

    • report.json, report.md, and other summary artifacts are projections over run_results.json and the manifest.

    • They must not redefine status semantics; they may only display or aggregate what run_results.json already encodes.

If manifest.json and run_results.json appear to disagree (e.g. artifacts exist but status is failed), run_results.json wins for execution truth. Such discrepancies should be treated as bugs or recovery cases, not new semantics.

6.1 Valid vs invalid combinations (summary)

Interpretations below rely on the output contract (docs/contracts/output-contract-v1.md) to define which artifacts are required for a module.

  • Valid:

    • failed + partial outputs → allowed; partial artifacts may exist but status remains failed.

    • failed + complete outputs → allowed; status failed still wins as execution truth.

  • Invalid (contract violation):

    • succeeded + missing required artifacts → invalid; either the outputs are missing or the status is wrong.

    • succeeded + only partial outputs where more are required by the output contract → invalid.

    • skipped or blocked + required outputs present → invalid; such outputs must be treated as stray or stale.

7. Group truth projection

For group runs:

  • Group-level run_results.json is the canonical persisted group-module rollup.

  • Per-member run_results.json files are the canonical member execution detail.

  • Read-side group consumers combine both through project_group_outcomes(...) using deterministic predicates such as:

    • any_member_usable

    • group_phase_terminal_failure

    • all_members_blocked

    • all_members_skipped

Group phase metadata may be persisted in additional files (e.g. aggregation_warnings.json), but group status must continue to be derived through the typed group outcome loaders and projection helpers.

8. Contract violation clause

  • Any state not representable by:

    • the allowed statuses listed in §2, and

    • the projection and precedence rules in §4–§7

  • is a contract violation and must be treated as an error:

    • Log a clear error.

    • Surface it in diagnostics.

    • Fail closed where appropriate.

Silent acceptance of such states is not allowed; they indicate either a bug or an incomplete migration and must be corrected.


9. Consumer rule

  • Consumers that need execution truth must always trust run_results.json over:

    • raw file presence/absence,

    • manifest.json,

    • or any other derived summary.

  • Any codepath that infers success/failure/skipped/blocked/requested/enabled without going through the typed run outcome loaders is a violation of this contract.


10. Contract violations (summary)

This section summarizes run outcome contract violations, how they are detected, and the expected behavior.

  • Invalid states (examples):

    • Status values outside the allowed set in §2.

    • Combinations of status fields and projection inputs that cannot be reconciled by the rules in §4–§7.

    • Codepaths that infer execution truth from file presence, manifest.json, or derived reports without consulting run_results.json.

  • Detection:

    • Typed loaders such as load_run_results(...) and load_run_outcome_context(...) that validate schema version and required fields.

    • Projection helpers in run_outcome_truth.py that assert invariants and raise on impossible combinations.

    • Tests that enforce precedence rules between run_results.json, manifest.json, and downstream reports.

  • Expected behavior:

    • Treat such states as fail-fast contract violations: surface clear diagnostics, fail closed where execution truth is ambiguous, and require code or data fixes before proceeding.