Risk OS Red Try Stoa

Stoa JSON Schema

This document describes the structure of stoa-registry.json, the JSON document produced by stoa scan.

Current schema version: 1.3

Versioning policy

version (1.01.1).

meaning) bump the major version (1.x2.0).

minor release.

for a given tree and configuration.

a scan that produces no AI (AI0xx) findings serializes byte-identically to 1.0 apart from schema_version — every 1.1 field below is emitted only when it carries data.

Schema 1.1 additions (v0.2)

On a finding (present only on AST/flow-based AI0xx findings):

FieldTypeMeaning
idstring"<rule_id>-<fingerprint[:12]>", the stable finding id
canonical_namestringe.g. STOA-LLM02-OUTPUT-EXEC (also the SARIF ruleId)
owaspobject{"llm_top10_v1_1": "LLM02", "llm_top10_2025": "LLM05"}
variantstringrule sub-variant (e.g. AI005 trust-remote-code)
flowarraytaint steps: `{role: source\propagation\sink, line, snippet}` (snippets redacted)
gate_eligiblebooltrue only for AI002 exec-class at high confidence
dimensionsarraydimension ids this finding contributes to
supersedesarrayrule ids this finding dedups (e.g. AI002/sql supersedes SEC003)
evidence_tagsarraye.g. system_role_interpolation, local_endpoint_observed

On an agent candidate: dimension_assessment — per-dimension exposure block: `{taxonomy: {id, version}, dimensions: [{id, group, assessability, exposure, score, contributing_findings, contributing_capabilities,

controls_observed, statement}]}. exposureelevated | moderate | low |

none-observed | not-assessable (never "safe"/"covered"); proxy-tier dimensions are capped at moderate. group (schema 1.3+) is a display grouping letter (e.g. AD on the default taxonomy) — empty string on a custom taxonomy that doesn't define groups. **On a finding:** dimensions — the dimension ids it contributes to. **Top-level:** dimension_summary — org rollup (per-dimension max exposure and agent counts); degraded_files` — files whose AST parse degraded.

Schema 1.2 additions (Assurance layer)

Three independent additions. Declared metadata and the contradiction detector are opt-in by presence: a scan with no stoa-declared.toml serializes byte-identically to 1.1 apart from schema_version. Autonomy inference is unconditional, like highest_severity — every agent candidate gets an autonomy_level, regardless of declarations. Permission tags (permission_tags on every agent candidate) are also unconditional, like capabilities.

On an agent candidate — declared metadata (present only when stoa-declared.toml declares this agent id): declared — the raw declared record: {name, owner, purpose, users, geography, production_status, autonomy_intent, data_classes, economic_authority}. economic_authority, when set, is {max_per_action?, daily_aggregate?, worst_case_customer_loss?}, each {amount: number, currency: string}.

On an agent candidate — autonomy inference (always present):

autonomy_level{level, signals, reason}. level ∈ `recommend_only |

human_approved | bounded_autonomous | unrestricted_autonomous | indeterminate — a static classification of how unattended the agent's side-effecting reach appears to be, derived from existing detectors (AI002 side-effecting sinks, AI003 approval-absence, a same-file bounding signal). signals — the evidence list, [{signal (a rule id or a named pattern like "approval_construct"/"bounding"), path, line}]. reason — populated only when level == "indeterminate"`: the classifier never guesses when signals don't cleanly resolve.

On an agent candidate — permission tags (always present, possibly empty): permission_tags — a higher-stakes layer on top of capabilities: move_funds, approve_transactions, sign_contracts, delete, communicate (an alias over email_send/messaging).

On a finding — the contradiction detector (DECL001-DECL007 only): declared_ref{path, key}, the declaration-side evidence (the stoa-declared.toml key path this finding contradicts), alongside the finding's own path/line (the code-side evidence). Cross-checks declared facts against what the scan actually observed — e.g. DECL001 fires when autonomy_intent is recommend_only/human_approved but the inferred autonomy_level is bounded_autonomous/unrestricted_autonomous. See docs/declarations.md for the full rule table.

Top-level: business{industries?, regulated_activities?, max_customer_dependency?, societal_risk_flags?}. governance{release_approval, incident_response, risk_acceptance?, harmful_output_policy?}. evidence — pointers only, grouped by category (testing, safety_testing, monitoring, contracts, vendor, historical, or any other caller-supplied category name), each entry {kind, ref, date?}. All three present only when stoa-declared.toml exists.

Schema 1.3 additions (AIUC-1 alignment)

Renames the default dimension taxonomy's ids and adds a display group field — see docs/dimensions.md for the full rationale and the old→new id mapping. This is the default taxonomy shipped with the binary (stoa-aiuc-8, v2.0, replacing stoa-default-8 v1.0); a custom taxonomy supplied via [dimensions] taxonomy is unaffected.

Two new optional declared fields, both additive and both attestation-only (never scored — see docs/dimensions.md):

business.societal_risk_flags (list, subset of `critical_infrastructure |

biosecurity_adjacent | mass_influence) and governance.harmful_output_policy (string pointer). Two new recognized evidence categories: safety_testing and vendor, same {kind, ref, date?}` shape as the existing categories.

assurance-packet/1.1 (produced by stoa export --assurance, not part of stoa-registry.json itself): the packet's 14 areas become 18, each now carrying a group key (one of index, AF, G) that groups them under AIUC-1's six standard categories plus a seventh, Stoa-only group for insurance-specific exposure. See docs/assurance-export.md.

Schema 1.4 additions (Runtime trace overlay)

All optional, emitted only by stoa runtime merge / stoa scan --with-runtime — a plain scan serializes byte-identically to 1.3 apart from schema_version. See docs/runtime.md.

On an agent candidate (merge only):

FieldTypeMeaning
runtime_evidenceobjectObserved-behavior summary for the analyzed window: {window: {start, end}, span_count, spans_by_kind, error_rate, observed_capabilities, observed_integrations, observed_providers, observed_models, capability_counts, integration_counts, high_impact_actions, high_impact_approved, approval_rate_high_impact, max_observed_amount, window_total_amounts, evidence_quality, delegations_to, trace_files}
liveness_statestringThe field reserved since 1.0, now live: "active" (spans observed in window) or "idle" (registry agent, zero spans). "deprecated" stays reserved — inferring it needs more than one window.

On a finding (RT-family only): trace_ref{file, line, span_id}, the trace-side evidence pointer, sibling to the DECL family's declared_ref. RT findings (RT001RT005) are appended by merge to the agent's findings list; the scan-time summary.findings counts are deliberately not rewritten (they describe the static scan) — merged RT findings are counted in the top-level runtime block instead.

On a dimension_assessment entry (merge only, the two proxy dimensions only, and only when spans cover the agent): assessability may become "runtime", accompanied by evidence_window ({start, end, span_count}, always non-empty — enforced by a property test) and runtime_basis (the observed signals the exposure re-bucketing used, so the bucket is auditable from the registry alone). Entries still labeled proxy remain capped at moderate exactly as before.

Top-level (merge only): runtime{analysis_schema, window, span_count, agents_covered, agents_total, unmatched_agents, rt_findings, evidence_quality}.

stoa diff ignores runtime_evidence, liveness_state, and RT-family findings unconditionally: a diff describes call sites added or removed in code, and runtime data varies run to run.

Companion schemas (separate documents, not part of the registry)

SchemaProducerNotes
stoa-trace/1.0the stoa.runtime SDKJSONL, one span per line; line 1 is a header record ({kind: "header", schema, sdk_version, redaction, dropped_spans}). Span fields: kind (`agent_run \llm_call \tool_call \action \approval \retrieval \delegation), trace_id, span_id, parent_span_id, agent_id (12-hex or null + agent_hint), start_ts/end_ts (ISO-8601 UTC — trace files are the one place timestamps live in content), status, redaction, and optional capability/integration/provider (scanner vocabulary ids; off-vocabulary values are flagged vocabulary: "custom"), model, tool, amount {amount, currency}, approval {approved_by, method}, approval_span_id, from_agent_id/to_agent_id (delegation), attrs (hashes/lengths by default — see redaction). Reserved fields: enforcement, session_id, cost`.
runtime-analysis/1.0stoa runtime analyzeDeterministic body given identical traces; wall-clock in header.generated_at only. agents (per-id summaries), unmatched_agents (never silently dropped), no_runtime_evidence (explicit).
runtime-baseline/1.0stoa runtime baselineCommitted like .stoa/approvals.toml, reviewed like code.
runtime-drift/1.0stoa runtime driftDrift events (`high \medium \info`) + the exact thresholds used.

assurance-packet/1.2 (produced by stoa export --assurance): the reserved observed status/provenance is live — Areas 12 (Monitoring) and 18 (Claims evidence) populate from a runtime-enriched registry, and RT findings join the contradictions table. A packet from a registry without runtime data is byte-identical to 1.1 apart from the schema string.

Schema 1.5 additions (Regulatory crosswalk)

A presentation/labeling layer that anchors findings to frameworks a reader already knows. It never participates in scoring — dimension scores, exposure buckets, and the proxy cap are byte-for-byte identical to 1.4 (guarded by a pre-change golden snapshot). Versioned as stoa-crosswalk-1 and overridable via stoa.toml [crosswalk] path.

On every finding: crosswalk{owasp_llm_2025, eu_ai_act, relation, so_what}. owasp_llm_2025 is one OWASP LLM Top 10 (2025) class (LLM01LLM10) or null where no honest class fits; eu_ai_act is one article (e.g. "Art. 15"); relation is exposure or control-observed; so_what is a plain-English gloss (says/never-says vocabulary — never "compliant"/"protected"/"secure").

On each dimension_summary dimension: crosswalk{owasp_llm_2025: [...], eu_ai_act: [...]}, the union of the framework tags of the rules that contributed to that dimension (OWASP codes sorted LLM01LLM10).

Top-level: crosswalk{id, version, owasp_llm_version, eu_ai_act_reference}, attributing the mapping to a reviewed version.

The pre-existing per-finding owasp object (schema 1.1) is unchanged and independent — the crosswalk lives under crosswalk, never overwriting it.

SARIF: results and rules gain owasp:<LLMxx> and euaiact:<article> tags alongside the existing stoa-dim:<dimension> tags. A blank OWASP mapping emits no owasp: tag rather than a fake one.

Top-level document

{
  "schema_version": "1.0",
  "tool": { "name": "stoa", "version": "0.1.0" },
  "repository": {
    "name": "payments-service",
    "root": ".",
    "git_ref": "abc1234",
    "base_ref": "origin/main"
  },
  "summary": { "...": "see below" },
  "agents": [ "...agent records..." ],
  "repository_findings": [ "...finding records..." ],
  "skipped_files": [ { "path": "node_modules/", "reason": "..." } ],
  "warnings": [ "...scan warnings, e.g. diff fail-open notices..." ]
}
FieldTypeNotes
schema_versionstring"<major>.<minor>"
tool.name / tool.versionstringProducer identity
repository.namestringSanitized (credentials stripped from remote URLs); falls back to the root directory name
repository.rootstringAlways "."; paths in the document are relative to it
repository.git_refstring \nullAbbreviated HEAD commit, when available
repository.base_refstring \nullThe --base ref, when diff-aware scanning was requested
agentsarrayAgent-candidate records, sorted by (path, symbol)
repository_findingsarrayFindings in files that are not agent candidates, sorted by (path, line, rule_id)
skipped_filesarraySkipped files or pruned directories (directory entries end with /) with reasons
warningsarray of stringsNon-fatal scan warnings (e.g. diff-gating fail-open)

summary

{
  "files_scanned": 347,
  "agent_candidates": 4,
  "high_confidence_candidates": 3,
  "integrations": 6,
  "findings": { "critical": 1, "high": 2, "medium": 5, "low": 0, "info": 3 },
  "new_findings": { "critical": 1, "high": 0, "medium": 0, "low": 0, "info": 0 },
  "suppressed_findings": 2
}

findings and new_findings count unsuppressed findings only. new_findings is all zeros unless a diff base was resolved.

Agent record

{
  "id": "9f2c41d0a3b7",
  "name": "refund_agent",
  "symbol": "refund_agent",
  "path": "src/refund_agent.py",
  "language": "python",
  "confidence": "high",
  "detection_score": 10,
  "evidence": [
    { "rule_id": "AGENT_LANGCHAIN", "line": 41, "description": "LangChain agent construct" }
  ],
  "providers": ["openai"],
  "frameworks": ["langchain"],
  "integrations": ["postgres", "stripe"],
  "capabilities": ["database_read", "payment_access", "tool_calling"],
  "call_sites": { "postgres": 1, "stripe": 2 },
  "last_touched_by": "Alice Smith",
  "last_commit": { "hash": "abc1234", "date": "2026-07-18T12:30:00-07:00" },
  "codeowners": ["@payments-team"],
  "findings": [ "...finding records for this candidate's file..." ],
  "highest_severity": "critical"
}

Notes:

source identity.

(see README). An agent record is always a candidate, never a confirmed agent.

not a runtime API call count.

email address). It is not ownership.

findings; deduplicate by fingerprint when aggregating.

findings.

Finding record

{
  "fingerprint": "3f7a9c2e51b8d4f0",
  "rule_id": "SEC001",
  "title": "Possible hardcoded API credential",
  "category": "secret",
  "severity": "critical",
  "confidence": "high",
  "path": "src/refund_agent.py",
  "line": 15,
  "column": 12,
  "snippet": "api_key = \"sk-pro…[REDACTED:a18c45f21a0e]\"",
  "remediation": "Load the credential from a secret manager or environment variable.",
  "suppressed": false,
  "suppression_reason": null,
  "is_new": true
}

Notes:

stable across pure line-number movement. Identical contexts in one file are disambiguated with an occurrence index.

in this document.

confidencelow | medium | high.

range of the diff against repository.base_ref; it is always false when no base was resolved.

Reserved field names

The following field names are reserved for future schema versions and must not be used for any other purpose by producers or consumers of this schema. They are not emitted in version 1.0 and carry no behavior today:

Reserved fieldFuture purpose
autonomy_level~~Reserved~~ — live since 1.2 (static autonomy inference)
loss_scenariosMapping of findings and capabilities to loss-scenario descriptors
liveness_state~~Reserved~~ — live since 1.4 (active/idle from the runtime overlay; deprecated still reserved)
policy_linesMapping to insurance policy-line identifiers
exposure_classNormalized exposure categorization

Reserving these names now prevents breaking schema changes later.