AI evaluation · AI safety · AI governance

The feature is the evidence chain.

FairMind sits between evaluator tools and governance decisions. It gives every result a subject, threshold, artifact, limitation, review state, finding, remediation, and rerun—so a test can become evidence without pretending to become compliance.

InputEvaluator output

Run a supported path or ingest a result from a tool your team already trusts.

FairMind addsChain of custody

Identity, threshold, artifacts, limitations, review, findings, and retests.

OutputDecision-ready package

Framework-pinned evidence for a stated cutoff and accountable reviewer.

Six jobs, one durable record.

The product is organized around the work required to defend a decision—not a collection of unrelated dashboards.

01 / System boundary

Define what is actually under review.

Organization, workspace, AI system, version, owner, deployment, intended use, and applicable frameworks.

02 / Evaluation run

Run or ingest a bounded test.

Record evaluator identity, metric version, prompt or dataset scope, threshold, result, and capability state.

03 / Evidence record

Preserve the artifact and its limits.

Bind provenance, integrity information, raw output, limitations, and the exact subject that produced it.

04 / Reviewer decision

Keep technical output separate from governance acceptance.

Candidate mappings require an accountable person to accept, reject, or narrow them with rationale.

05 / Finding and retest

Retain failure history after remediation.

Link finding, risk, owner, remediation, exception, and clean rerun without erasing the earlier failure.

06 / Assurance package

Compile decision-ready evidence.

Export a framework-pinned index for a stated evidence cutoff—decision support, never certification.

The Evidence Passport is the target exchange contract.

FairMind’s novelty is not another universal trust score. It is a portable, reviewable unit that binds technical evidence to the human decisions and later changes that give it governance meaning.

01Exact subject and version
02Evaluator and metric identity
03Pre-registered threshold
04Artifact integrity and raw output
05Limitations and capability state
06Reviewer decision and rationale
07Finding, remediation, and exception
08Superseding rerun without erased history

Capability state travels with the evidence.

These are implementation boundaries, not maturity grades. A route can be useful and still be bounded.

Evaluator pathStateWhat it can produceBoundary
Environmental evidencevalidatedEnergy, carbon, water, uncertainty, mitigation, exception, and reviewer records.Real hardware evidence is still required before paper claims are treated as observed.
Provider-backed LLM judgeexternal_providerBounded assessment with provider identity and rubric version.Requires a configured provider and preserves that dependency as a limitation.
AI BOM fairness profilemetadata_onlyTransforms inventory, bias, remediation, and unknown metadata.It structures supplied metadata; it does not independently evaluate the system.
Multimodal group analysismetadata_onlyDistribution analysis over caller-supplied labels.Media content itself is not evaluated.
Tabular biasunavailableNo evidence-grade active route.Unavailable must remain unavailable until a real contract closes.
Modern LLM biasunavailableMock output cannot enter assurance packages.Browser-generated results and pass fallbacks are excluded.

Start with one real evaluator output

Make the evidence boundary visible.