Security teams do not need another assistant that produces polished prose while quietly detaching itself from the evidence, permissions, and workflow state behind the question. They need a reasoning layer that can help with difficult analysis while preserving the difference between what was observed, what a deterministic service established, what a model inferred, what is still missing, and what an authorized person actually decided.

That is the operating idea behind FORGE by Threat Foundry. Forge brings evidence-grounded analysis into investigation, triage, cases, detection design, validation, tuning, and managed operations. It works inside existing product workspaces, uses a closed set of bounded tools, and keeps the source and decision record attached to every recommendation.

Forge proposes and explains from authorized evidence. Deterministic services verify. Authorized humans decide.

Why evidence-grounded reasoning matters

A plausible answer is not the same as an operationally safe answer. In security operations, a useful recommendation depends on questions that ordinary chat does not reliably answer:

  • Is the evidence from the correct tenant, user, case, finding, detection, and revision?
  • Is the source current, authorized, sufficiently complete, and permitted for this workflow?
  • Which facts were directly observed, and which conclusions came from deterministic product logic?
  • What did the model infer, and what evidence is missing?
  • What action is the analyst allowed to take, and which decisions require independent review?

Forge treats those questions as part of the capability, not as fine print. Tenant, identity, permissions, workspace, source revision, policy, retrieved evidence, deadlines, cancellation state, and citations are checked throughout a run. If a material dependency changes, prior analysis becomes stale or stops instead of silently continuing with yesterday’s authority.

The Forge workflow: context, plan, evidence, analysis, decision

A Forge interaction is more structured than an open-ended prompt. The analyst begins in an existing workspace and asks a bounded question. Forge then builds a concise plan: what capability is being used, which evidence categories are required, which closed tools may be called, what is missing, and which limitations will remain.

Server-owned tools retrieve only the current context allowed for that tenant and user. Exact references are resolved first. When semantic retrieval is enabled, healthy, tenant-scoped, and authorized, it can supplement exact evidence; similarity remains advisory and never establishes permission, provenance, or proof. Retrieved text is treated as untrusted material and is kept separate from server-owned source metadata.

The final view should let the analyst inspect:

  1. Authoritative state: the current lifecycle, ownership, severity, approval, deployment, validation, or policy state established by the product.
  2. Observed facts: bounded evidence the authorized tools actually returned.
  3. Forge inferences: model-supported interpretation that must not be confused with authoritative state.
  4. Missing evidence: the unanswered questions or unavailable sources that constrain the conclusion.
  5. Recommendations: bounded next steps, limitations, and the human review required before anything material changes.

This structure makes a recommendation easier to challenge. A reviewer can ask whether the evidence supports it, whether the deterministic state contradicts it, and whether the proposed next step stays within the current authority.

Closed tools keep capability and authority legible

Forge does not receive a general-purpose computer. Its production tool registry is finite and server owned. Each tool has a fixed name, version, purpose, action class, permission set, argument schema, result schema, timeout, result bound, and invocation limit. Model output cannot add a tool, enlarge a limit, choose another tenant, supply credentials, select arbitrary URLs, or override the execution context.

The closed-tool model deliberately excludes arbitrary shell, Python, SQL, file-system, browser, network, credential, provider-administration, approval, deployment, and response execution. This does not make every recommendation correct. It makes the operational boundary inspectable and enforceable.

Forge also keeps a safe timeline of planning, tool authorization, retrieval state, cancellation, supersession, recommendation persistence, and human decisions. It does not need to store raw credentials, provider authorization headers, vectors, hidden reasoning, or unrestricted provider payloads to make the workflow reviewable.

Use Forge for investigation and triage

In a case or triage workspace, Forge can help organize current evidence into an investigation story. It can surface observed and missing evidence, recurring observations, related authorized context, unresolved questions, likely next steps, and a bounded stakeholder-update draft.

Deterministic state keeps precedence. Forge cannot decide that a finding has a lower severity than a policy floor, claim containment or impact that the case does not establish, change a disposition, link or close a case, confirm an incident, or perform response. Similarity may suggest related evidence, but it cannot prove that two cases are the same event.

When follow-up work is appropriate, Forge may propose a small set of bounded case-task drafts. The exact proposal is bound to the current tenant, user, run, evidence, and confirmation. An authorized local user must review it. The tasks remain unowned, undated, non-blocking drafts; they do not change the case lifecycle or assign operational authority.

Use Forge for detection and validation analysis

Detection engineering benefits from explanation, but it also contains several states that must remain deterministic. Forge can analyze ATT&CK coverage, design a detection or a detection chain, plan validation, and interpret stored validation evidence. Closed read tools can retrieve bounded definitions, candidates, telemetry and mapping context, coverage state, test metadata, Assurance state, health, and exact stored outcomes.

The distinction between proposal and proof is critical:

  • A generated detection is a candidate, not an approved control.
  • A validation plan is not an executed test.
  • A model interpretation cannot replace the exact deterministic validation outcome and reason.
  • A related ATT&CK or D3FEND mapping is context, not proof of coverage or effectiveness.
  • A draft created through Apply-as-Draft is not active, deployed, validated, or approved.

Apply-as-Draft requires current local authority, a substantive rationale, fresh source and policy state, current citations, and deterministic validation. Forge cannot approve, activate, deploy, schedule, export, launch validation, set the outcome, contact or mutate a provider, or perform a response action.

Use Forge to prepare governed tuning work

Tuning is where apparently helpful automation can weaken coverage without making the loss obvious. Forge therefore separates candidate classification, tuning analysis, proposal preparation, independent review, retesting, approval, and deployment.

Before bounded preparation is eligible, deterministic services evaluate current source and fingerprints, evidence sufficiency and freshness, confidence, materiality, protected state, allowed tuning type, positive and negative controls, current permissions, citations, cooldowns, and open-candidate limits. The more restrictive result wins. Unknown materiality, stale evidence, a deterministic fallback, protected state, unsupported tuning, or missing controls blocks automatic preparation.

When all gates pass, Forge may prepare one current analysis, one review-ready proposal, an optional unapproved draft candidate, and exact references to retest work. It does not accept the proposal, approve the candidate, run provider-backed validation, modify an active detection, write to the provider, or deploy. Independent reviewers still decide whether the proposal should move forward, and current retest evidence still determines readiness.

Use Forge Managed Operations without centralizing evidence

MSSP and MSP teams need portfolio guidance, but copying raw customer evidence into a central service plane creates the wrong trust boundary. Forge Managed Operations keeps retrieval and model contact inside the customer tenant. An authorized Control Center operator selects one closed analysis purpose for an eligible work item; the tenant revalidates the source, entitlement, workspace, and current authority before using its own Forge runtime.

Control Center receives an evidence-free projection: analysis state and currentness, advisory priority, evidence sufficiency, recommended workflow, limitations, and deterministic restrictions. Operational priority, SLA, assignment, escalation, suppression, case state, triage state, detection state, validation state, and tuning state remain authoritative outside Forge.

If deeper review is required, an authorized operator can use a short-lived, read-only tenant handoff to inspect the current tenant-owned analysis. That path does not create cross-tenant retrieval, model analysis across customers, automated assignment, priority changes, proposal acceptance, provider write-back, or response authority.

Design for degraded states, not just ideal runs

External models can be unavailable, rate limited, misconfigured, or unable to return valid structured output. Semantic retrieval can also be disabled or unavailable. Forge makes those states visible and preserves the deterministic product workflow.

Exact structured evidence can remain usable without semantic search. A deterministic fallback remains labeled as a fallback. No proposal or draft is created when a required external analysis fails. An inactive provider configuration is not used as an undisclosed second route. If the active BYOAI provider or model changes after review, the prior handoff does not silently redirect.

This behavior matters because operational resilience is not the absence of failure. It is the ability to fail without inventing evidence, mislabeling the source of a conclusion, or taking authority the system does not have.

A practical operating playbook

  1. Start with one decision. Pick a case question, triage finding, coverage gap, validation result, or tuning candidate that already has an accountable owner.
  2. Inspect the source state first. Confirm tenant, revision, lifecycle, permissions, evidence age, provider route, and the deterministic facts that must not change.
  3. Ask a bounded question. State the purpose and desired decision, not an instruction to take an uncontrolled action.
  4. Review the Forge Plan. Check the required evidence, closed tools, missing evidence, limits, and whether synthesis can support the question.
  5. Challenge the analysis. Separate observed facts from inferences and confirm that citations point to current authorized evidence.
  6. Choose the next human-controlled step. Create or review a draft, collect missing evidence, run an approved test, assign work, or decline the recommendation.
  7. Measure the outcome. Track whether the workflow reduced investigation time, improved evidence completeness, produced useful validated work, or prevented an unsafe shortcut.

Measure useful assistance without rewarding automation theater

Good Forge metrics should describe the quality and safety of the operating loop: current versus stale analyses, sufficient versus insufficient evidence, recommendations accepted or revised by humans, deterministic overrides, prepared drafts that passed independent review, retest completion, time to the next accountable decision, and workflows that continued safely during provider degradation.

Avoid measuring success as the number of prompts, model responses, generated rules, or automatically created items. Volume can increase while decision quality falls. The stronger question is whether Forge helped the team reach a better-supported decision with less avoidable toil while keeping authority, provenance, and limitations visible.

That is the role Forge is designed to play: a reasoning layer inside governed security operations, not an autonomous replacement for the controls, evidence, and people that make a security decision defensible.