Risk OS Red Try Stoa

Dimension exposure

Every agent candidate is assessed across a taxonomy of risk dimensions and rendered as a Dimension Exposure Matrix at the top of the HTML report. The default taxonomy has eight dimensions grouped under six categories — Data & Privacy, Security, Safety, Reliability, Accountability, Society. Five dimensions are assessable statically, three are proxy signals flagged for runtime follow-up — all with line-level evidence.

The table below covers what static analysis can score. The parts that need governance documentation or third-party testing evidence (most of Safety, Accountability, and Society) live in stoa-declared.toml and surface in stoa export --assurance instead, under the same six categories plus a seventh, Stoa-only category for insurance-specific exposure.

The eight default dimensions

GroupDimensionAssessabilityWhat static analysis sees
A — Data & PrivacyBoundary leakagestrongsensitive data leaving via model calls or egress
B — SecurityMandate overreachstrongreach beyond declared scope
B — SecurityInjection & tamper surfacepartialprompt-injection / supply-chain surface (robustness is runtime)
B — SecurityControl coverage gappartialauth / validation / rate-limit / observability
C — SafetyUnreviewed high-impact actionstronghigh-impact actions without an observed approval
D — ReliabilityOutput fidelitypartialunsafe model-output handling (correctness is runtime)
D — ReliabilityConduct variabilityproxyonly config signals (e.g. unpinned sampling)
D — ReliabilityDependency driftproxyonly upstream-pin signals

Notice groups E (Accountability) and F (Society) have no scanned dimension — that's deliberate, not a gap to be padded. A static code scan has no way to assess vendor due diligence or societal-scale misuse risk; those live entirely in the declared/ingested layers of the assurance packet.

Assessability tiers cap what Stoa may claim. A proxy dimension can never render elevated — it is capped at moderate, enforced by a property test. Stoa must never imply it measured behavior it only saw a config signal for.

The runtime trace overlay adds a fourth tier, runtime: when trace evidence covers an agent's window, the two proxy dimensions re-bucket from observed signals for that agent, for that window — no longer capped, in either direction — always carrying an evidence_window and the observed basis. Entries without runtime evidence stay proxy and stay capped.

Scoring (deterministic)

Per agent, per dimension:

score = min(100,
    Σ finding_weight(severity, confidence)   # e.g. critical×high-conf = 40
  + Σ capability_weight                       # each mapped capability contributes
  − Σ control_credit)                          # observed controls subtract (floor 0)

Buckets: 0 → none-observed · 1–24 → low · 25–54 → moderate · ≥55 → elevated, then the proxy cap applies. Weights live in data/dimensions.toml — changing them bumps the taxonomy version, so score changes are always attributable to a code change or a declared taxonomy change, never a silent recalibration.

Observed controls (approval, authentication, validation, rate-limit, observability, deterministic sampling, pinned model) subtract exposure — the one place Stoa reports good news, always phrased as "observed".

Suppressed findings contribute zero but remain listed in the drill-down.

Exposure values

elevated | moderate | low | none-observed | not-assessable. Never "safe", "covered", or "compliant".

Custom taxonomies

# stoa.toml
[dimensions]
taxonomy = ".stoa/dimensions.toml"   # replaces the default

A custom file declares its own [[dimensions]] and [rule_dimensions] / [capability_dimensions] maps. Any rule left unmapped falls into a reserved unclassified dimension that always renders — a custom taxonomy cannot silently drop findings from the dimensional view. The taxonomy id+version is embedded in every registry, so stoa diff across mismatched taxonomies exits 2 rather than producing a misleading comparison.

Flags: --no-dimensions (skip assessment + matrix), --taxonomy PATH.

Machine interface

The registry's per-agent dimension_assessment block and the top-level dimension_summary are the machine interface: "read stoa-registry.json and address all elevated boundary-leakage contributors" is a valid agent instruction with zero extra tooling. SARIF results carry a stoa-dim:<dimension> tag so GitHub Code Scanning can filter by dimension. Every dimension entry also carries group (one of AD, or empty for a custom taxonomy that doesn't use groups).

What Stoa says / never says

Stoa saysStoa never says
"Exposure observed" / "none observed""Covered" / "protected" / "compliant"
"Proxy signals only — runtime evaluation required""Behaviorally stable" / "drift-free"
"Controls observed: interrupt gate""Risk mitigated"
"Assessed across 8 dimensions: 5 direct, 3 proxy""Full coverage across 8 risk dimensions"