Run a supported path or ingest a result from a tool your team already trusts.
AI evaluation · AI safety · AI governance
The feature is the evidence chain.
FairMind sits between evaluator tools and governance decisions. It gives every result a subject, threshold, artifact, limitation, review state, finding, remediation, and rerun—so a test can become evidence without pretending to become compliance.
Identity, threshold, artifacts, limitations, review, findings, and retests.
Framework-pinned evidence for a stated cutoff and accountable reviewer.
Six jobs, one durable record.
The product is organized around the work required to defend a decision—not a collection of unrelated dashboards.
Define what is actually under review.
Organization, workspace, AI system, version, owner, deployment, intended use, and applicable frameworks.
Run or ingest a bounded test.
Record evaluator identity, metric version, prompt or dataset scope, threshold, result, and capability state.
Preserve the artifact and its limits.
Bind provenance, integrity information, raw output, limitations, and the exact subject that produced it.
Keep technical output separate from governance acceptance.
Candidate mappings require an accountable person to accept, reject, or narrow them with rationale.
Retain failure history after remediation.
Link finding, risk, owner, remediation, exception, and clean rerun without erasing the earlier failure.
Compile decision-ready evidence.
Export a framework-pinned index for a stated evidence cutoff—decision support, never certification.
The Evidence Passport is the target exchange contract.
FairMind’s novelty is not another universal trust score. It is a portable, reviewable unit that binds technical evidence to the human decisions and later changes that give it governance meaning.
Capability state travels with the evidence.
These are implementation boundaries, not maturity grades. A route can be useful and still be bounded.
| Evaluator path | State | What it can produce | Boundary |
|---|---|---|---|
| Environmental evidence | validated | Energy, carbon, water, uncertainty, mitigation, exception, and reviewer records. | Real hardware evidence is still required before paper claims are treated as observed. |
| Provider-backed LLM judge | external_provider | Bounded assessment with provider identity and rubric version. | Requires a configured provider and preserves that dependency as a limitation. |
| AI BOM fairness profile | metadata_only | Transforms inventory, bias, remediation, and unknown metadata. | It structures supplied metadata; it does not independently evaluate the system. |
| Multimodal group analysis | metadata_only | Distribution analysis over caller-supplied labels. | Media content itself is not evaluated. |
| Tabular bias | unavailable | No evidence-grade active route. | Unavailable must remain unavailable until a real contract closes. |
| Modern LLM bias | unavailable | Mock output cannot enter assurance packages. | Browser-generated results and pass fallbacks are excluded. |
Start with one real evaluator output