Metrics, suites, scans, artifacts, and reruns.
Compare the job, not the logo wall
Evaluators test. Governance manages. FairMind binds evidence.
FairMind is most useful between the tools that produce technical results and the processes that need defensible, reviewable evidence. It complements both categories instead of pretending to replace them.
Subject, threshold, limitation, review, finding, and retest.
Frameworks, policies, approvals, reports, and accountability.
The missing layer is evidence continuity.
This category comparison is deliberately vendor-neutral. Products vary; the handoff failure is consistent.
| Concern | Evaluator tools | Governance platforms | FairMind evidence layer |
|---|---|---|---|
| Primary job | Run a technical evaluation | Manage program, policy, or inventory | Preserve evaluator output through review and decision |
| System identity | Often local to the run | Usually an inventory record | Bound to the evidence subject, version, owner, and deployment |
| Threshold | Metric or suite configuration | May live in policy text | Pre-registered and stored beside the result |
| Raw artifact | Native output | Often linked or summarized | Indexed with provenance and integrity information |
| Framework mapping | Usually outside scope | Often manual or workflow-specific | Candidate mapping with independent reviewer decision |
| Failure history | Reruns may replace the working view | Tracked as issue or action | Finding, remediation, exception, and superseding retest stay linked |
| Final claim | Technical result | Governance program state | Decision support; never automatic compliance or certification |
Bring the evaluator you already trust.
The roadmap prioritizes artifact adapters for Giskard, MLflow, Fairlearn/AIF360, ModelScan, and garak/PyRIT before adding more evaluator engines. Those are adapter targets, not claimed live integrations.
The evaluator is not the governance decision.
Keep your current test stack