AI evaluation · AI safety · AI governance
Bring your tests.Produce evidence you can defend.
FairMind is the vendor-neutral evidence workbench between your AI test tools and your governance process. It binds every result to the system, threshold, artifact, limitation, reviewer, finding, remediation, and retest that give it meaning.
- Input
- Your evaluator output
- FairMind adds
- Context + chain of custody
- Decision
- Always human
Evaluator output
validated path
Environmental assessment
- Subject
- Claims Assistant 2.4.1
- Evaluator
- FairMind-E / environmental
- Threshold
- ENV gate / bounded
- Artifact
- energy + carbon + uncertainty
Real hardware measurement is still required before research claims are treated as observed.
Accountable workflow
Watch FairMind preserve the evidence boundary as the same record moves into review.
FairMind connects evaluators to accountable governance. It does not replace the evaluator, infer compliance, issue certification, or approve an AI system because a dashboard turned green.
See the evidence graphThree teams. One evidence record.
Engineers, reviewers, and assurance owners enter at different points. FairMind keeps them anchored to the same AI system and the same source artifacts.
For ML engineers and security testers
Keep the result attached to the thing you actually tested.
Register the system boundary, bring or run an evaluator, record the threshold before interpretation, and preserve the source artifact with its limitations.
- SystemName the exact subject and version
- EvaluationRecord scope, threshold, and result
- EvidencePreserve artifact identity and limitations
Output Reviewable evidence record
For model-risk, compliance, and independent reviewers
Review what the evidence proves—and where it stops.
Inspect capability state, provenance, thresholds, freshness, and limitations. Accept, reject, or narrow candidate control mappings with your rationale attached.
- TruthSee validated, bounded, or unavailable
- MappingReview the proposed control relationship
- DecisionRecord reviewer identity and rationale
Output Human-reviewed evidence mapping
For AI assurance leads and accountable owners
Make release decisions with failures, fixes, and unknowns intact.
Turn evidence gaps into findings, assign remediation, require a linked retest, pin the framework version, and export the evidence index behind the decision.
- FindingOwn the blocker and remediation
- RetestLink the clean run to the failure
- PackageExport decisions, limits, and history
Output Decision-ready assurance package
The graph is the product.
Not another score. A durable relationship between the test, the evidence, the reviewer, and what happened next.
Focus or select a node to inspect why it exists.
The exact model, agent, application, version, deployment, owner, and intended use.
Creates the subject boundary before any result is interpreted.Capability truth is part of the interface.
Current implementation state is shown per evaluator. It describes what evidence a route can produce today—not compliance, certification, or product vision.
Energy, carbon, water, and uncertainty records
Strongest evidence-grade vertical currently in this repository.
Provider-evaluated response assessment
Requires a configured external model provider and recorded provider identity.
Media metadata and statistical summary
The current path does not evaluate image, audio, or video content.
Software and model inventory profile
Transforms supplied metadata; it does not independently verify the subject.
No defensible artifact from the current route
The interface calls an evaluator endpoint that is not present.
No defensible evidence artifact
The current interface produces mock results and must not be exported as evidence.
Workbench records are still mutable and are not yet the immutable Evidence Passports defined by the target exchange contract.
Inspect the recovery roadmapUse the evaluators you already trust.
FairMind’s opening is not to rebuild every test engine. It is to preserve outputs from model, data, LLM, agent, security, provenance, and environmental tools in one reviewable evidence chain.
Environmental evidence · provider-backed LLM judge · AI BOM transformation
Giskard · MLflow · Fairlearn / AIF360 · ModelScan · garak / PyRIT
System identity · thresholds · limitations · review · remediation · framework reuse
Research ships with its limitations attached.
Methods, artifacts, measurement status, and what remains unproven belong beside the work—not in a footnote after the claim.
Environmental evidence with provenance and uncertainty
Energy, carbon, water, mitigation, exception, and reviewer records. Smoke runs prove the path; real hardware evidence is still required for paper claims.
See research statusA fairness evidence profile for AI inventory records
Transforms supplied BOM, bias, remediation, and unknown metadata into a reviewer-facing profile. Metadata remains metadata—not independent verification.
Inspect the sourceStart with the test output you already have.