Skip to main content

AI You Can Audit, Not Just Trust

Every Stack9 AI Agent runs under a documented AI Policy, a named owner, an evaluation gate before release, and a full audit record — the same governance discipline April9 applies to information security, extended to AI rather than layered on top of it.

Design Position

No Training or Fine-Tuning on Client Data

Stack9 AI Agents are never trained or fine-tuned on client data. Client content is what an agent reads at the moment a question is asked, not what shapes a model's weights.

A fine-tuned model absorbs content into weights that cannot be cited, version-stamped against a source change, or corrected without retraining. Grounding agents on live, version-stamped content instead means every answer reflects the source as it currently stands, and cites the version it used.

Assurance

Named Risks, Each With a Control

Governance sits inside April9's ISO 27001 Information Security Management System, held for over five years. A documented AI Policy applies to all personnel on stated principles — fairness, transparency, accountability and human-in-the-loop oversight — with a named owner for every AI system in scope, and AI awareness training delivered to all employees. Every change to an agent is raised as a Request for Change, scored against the client-approved baseline, and approved by the client's named approver. Decision authority stays with the client — always.

Carried on the project risk register, reviewed weekly during delivery and in the monthly report once in service.
RiskControl
Wrong or unsupported answerAgents answer only from client-approved authoritative sources and decline and escalate where those sources do not support an answer. A separate verification agent checks the assembled answer against what was retrieved.
Drift after a source changeBaseline evaluation sets are re-run on every content refresh, so the effect on accuracy is measured before release rather than discovered after it.
Unannounced model changeThe foundation model is pinned to a version; a new version is regression-scored before adoption, and the client is never moved silently. No model is mandated.
Inequitable serviceConsistency across phrasing, location and category is a release gate, scored on the same evaluation harness as accuracy.
Data leakagePer-client isolated environment, no training path for client data, and all AI processing stays inside the Australian region (AWS ap-southeast-2, Sydney).
Over-reliance by staffOverride and disagreement rates on AI drafts are tracked as a standing signal — staff ceasing to override is treated as a signal to investigate, not as success.
Evaluation

Evaluation as a Release Gate

Success is defined before an agent is configured, as a test harness — not a judgement made under schedule pressure once one is already live.

The client approves the evaluation set and pass threshold before the first agent is configured. A version scoring below that threshold is not promoted — on every change, and every content refresh.

Scoring runs on Amazon Bedrock AgentCore Evaluations, the same tooling used in Stack9 AI Studio, and checks four things: whether the outcome is correct; whether the right source is cited at the right version; whether the agent correctly declines and escalates where content does not support an answer; and whether a materially equivalent question reaches the same outcome. Whole conversations are tested too — a synthetic user with a persona and a goal holds a full conversation, scored by an AI judge against the expected outcome.

Bias & Harm

Equity and Harm, Tested Not Assumed

Nothing About the Person

No demographic attribute reaches any agent, and no one is scored, ranked, approved or refused by a model. A materially different answer to an equivalent question is treated as a defect, not a variation.

Three Named Harms

A person misled by an unsourced answer, a person in distress met with a scoped and guarded response, and a decision trusted to AI too far — findings are Complies, Does Not Comply or Attention Required.

Agents run under AgentCore Guardrails blocking hate, harassment, self-harm, violent or sexual content, manipulation and prompt-injection attempts, regardless of framing. Where a session shows signs of crisis, the agent stops engaging on the substance, surfaces a crisis-support reference, and hands off to the client's contact channel flagged as urgent.

Explainability

What Each Audience Sees

AudienceWhat They NeedHow Stack9 Delivers It
End userThe source behind every answerEvery answer links the source it used, version-stamped, plus what was retrieved, from which service, and when live data was queried.
Back-office userTo agree with or override a finding, not accept a whole documentA structured chain — provisions applied, data matched, and the finding — shown against the source, with the agent's reasoning open on request.
Administrators & assurance staffThe full execution record behind any past answerEvery request, response, tool call and identity used, plus the agent version, prompt, sources and model that produced it — versioned and change-controlled.

The whole interaction can be replayed step by step in Stack9 AI Studio.

Auditability

Every Interaction Captured as a Record

Interaction Record

Every prompt and response, each tool call, the identity it ran under, and its result — plus sources retrieved with their version and any refusal or escalation — captured as a record, not a log line.

Three Decision Layers

The agent's (which skill, which sources, where it declined), the person's (what a back-office user accepted, edited or overrode), and the configuration's (each RFC, its evaluation result and named approver).

Configuration in Force

Agent version, system prompt, skills, tool and source grants are retained as versioned artefacts alongside the interactions they produced — so a past answer is examined against the instructions actually in force.

Export & SIEM

Retained in the Australian hosting environment and exported in open formats, or streamed to the client's own SIEM — the client's choice at design, mapped to National Archives of Australia principles for AI-generated records.

Lifecycle

Five Loops, No Retraining in Any of Them

Content

Document-based content is re-indexed and version-stamped on schedule and on change; a source change triggers a baseline re-run.

Configuration

The versioned agent configuration is the unit of change — every change is an RFC, re-scored before release.

Model

The foundation model is pinned to a version; a new version is regression-scored before adoption.

Monitoring

Per-agent telemetry on accuracy, citation coverage, refusal rates, latency and cost — reported for the agent, not blended platform-wide.

Feedback

Staff accept, edit and override patterns feed the configuration loop; user ratings are written against the specific interaction.

On the Roadmap

ISO 42001 — Targeted for Q2 2027

AI management-system certification is targeted for Q2 2027, extending the certified-governance approach already applied under April9's ISO 27001 ISMS to the AI lifecycle. Not yet held.

Stated Plainly

What This Governance Does Not Claim

  • ISO 42001 is a target, not a certification held today.

  • No external AI-ethics reviewer is claimed. Governance rests on April9's AI Policy under the ISMS, client-approved evaluation sets and change-controlled releases.

  • Fairness cohorts, scenario sets and thresholds are set with each client during design. There is no universal pre-built benchmark.

  • An AI-generated answer is bounded by model inference and retrieval, so it can take longer than a conventional transaction. Responses stream, and AI timings are reported separately.

  • Capturing every prompt and response over a multi-year retention period is a real, compounding storage cost. It is priced in, not absorbed silently.

Governance You Can Verify

If you have a security questionnaire, an AI governance framework or a jurisdiction-specific policy to map against, that mapping is part of discovery.