What MO§ES™ benchmarks.
Operator evals, performative benchmarks, bespoke enterprise evals, and the full module set. Each module answers a specific enterprise question with declared data requirements and outputs.
See What We Benchmark Run a 30-Day Operator EvalEvaluate the humans operating AI.
What is it?
Measures how a person or team actually operates AI systems across tasks, workflows, models, and operating conditions. Examines observable operating behavior — not AI knowledge tests or self-reported proficiency.
What data is required?
Canonical operator telemetry: token counts (input, output, cache read, cache write). Level 2 adds timestamps, sessions, model/tool identifiers, acceptance signals, retries, errors. No prompt content required.
What does it produce?
Per-operator performance profile: leverage, yield, construction, consistency, iteration structure, field position, archetype, volatility, top-performer status.
What enterprise question does it answer?
How do our people actually operate AI — and who is performing best?
Benchmark the interaction, not just the model.
What is it?
Compares how operators actually perform under observed or defined operating conditions. Not self-reported proficiency, certification, or generic knowledge tests. Benchmarks may compare operator vs operator, operator vs cohort, team vs team, before vs after, cohort vs reference field, and model × operator pairings.
What data is required?
Operator eval results from at least one measurement window. External benchmarks require reference population data (SigRank field). Internal benchmarks require cohort data from multiple operators or teams.
What does it produce?
Percentile bands (median, top 25%, top 10%, top 5%, top 1%), field position, cohort distribution, team comparison, model/operator fit scores, before/after deltas.
What enterprise question does it answer?
How do our operators compare — to each other, to external reference, and over time?
Your company should not inherit someone else's definition of AI proficiency.
What is it?
Company-specific evals built around your workflows, roles, models, tasks, and operating conditions. Enterprises define roles, workflows, tasks, models, systems, desired behaviors, constraints, operational goals, risk boundaries, and performance questions. MO§ES™ turns those into repeatable evals.
What data is required?
Workflow definitions, role definitions, model inventory, task taxonomy, and the enterprise's performance questions. Configurable via CLI, TUI, and MCP. Saveable as JSON.
What does it produce?
A company-specific eval framework you own. Reusable across future pilots and continuous measurement. Carries governance metadata, provenance, and decision-use labels.
What enterprise question does it answer?
How do our people perform on the work that actually matters to us?
The output of the eval.
What is it?
Structural analysis of operator behavior: field position, distribution, longitudinal movement, consistency, volatility, capability concentration, operating structure, archetypes, metric divergence, distance from reference populations, band stability, cohort shape, top performers, underdeveloped capability, team composition, operator similarity, AI learning curves.
What data is required?
Operator eval results across at least one measurement window. Longitudinal analysis requires 2+ windows.
What does it produce?
Operator profiles, cohort distributions, archetype assignments, band placements, trajectory estimates, divergence flags.
What enterprise question does it answer?
What does our AI operating population actually look like — structurally?
Map capability across the organization.
What is it?
Mapping AI capability across people × teams × workflows × models × systems. Shows where capability exists, where it is absent, and where dependencies or bottlenecks form.
What data is required?
Operator eval results + organizational structure (teams, departments, reporting lines) + workflow definitions + model inventory.
What does it produce?
Capability topology map, concentration analysis, dependency map, bottleneck identification, department comparison.
What enterprise question does it answer?
Where is AI capability concentrated, and where are the structural risks?
Which operators fit which workflows?
What is it?
Analysis of how well operators, models, and tools fit specific workflow stages. Identifies friction points, capability gaps, and repeated failure patterns.
What data is required?
Workflow stage definitions + operator eval results + model/tool usage data.
What does it produce?
Per-workflow operator fit scores, model/workflow fit scores, friction point maps, capability gap analysis.
What enterprise question does it answer?
Is the constraint the operator, the tool, or the workflow?
Which models amplify which operators?
What is it?
Analysis of operator × model pairings. Identifies which models amplify which operators, which pairings break down, and when changing the model helps vs when operator development is the better intervention.
What data is required?
Operator eval results with model identifiers + multi-model deployment data.
What does it produce?
Model/operator fit matrix, pairing performance scores, model sensitivity analysis, intervention recommendations (model change vs operator development).
What enterprise question does it answer?
Best model for the task. Best operator for the task.
How is capability distributed across teams?
What is it?
Analysis of how capability is distributed across teams using the same AI stack. Identifies operator complementarity, similarity clusters, workflow coverage, and dependency risk.
What data is required?
Operator eval results + team membership + workflow assignments.
What does it produce?
Team capability distributions, complementarity scores, similarity clusters, coverage maps, dependency risk flags.
What enterprise question does it answer?
Why do teams using the same AI stack operate differently?
Where is capability concentrated in too few people?
What is it?
Identifies whether advanced AI capability is organizational or isolated in a few individuals. Flags concentration risk — what happens if your top 1% leaves?
What data is required?
Operator eval results across the cohort + role/team assignments.
What does it produce?
Concentration indices, dependency maps, risk flags, succession recommendations.
What enterprise question does it answer?
Do we have organizational capability or a few power users?
Which operators operate AI the same way?
What is it?
Clustering analysis that identifies operators with similar operating behavior. Reveals archetypes, peer groups, and hidden capability that volume metrics miss.
What data is required?
Operator eval results across multiple metrics. (EVAL-014: in development)
What does it produce?
Archetype assignments, similarity clusters, peer group identification, hidden-capability flags.
What enterprise question does it answer?
Which operators operate AI the same way — and who are the hidden high-performers?
How fast are operators improving?
What is it?
Longitudinal analysis of operator improvement over time. Identifies acceleration, plateau, reversion, and emergence of new operating behavior.
What data is required?
Operator eval results across 2+ measurement windows. 3+ recommended for stable trajectory estimates.
What does it produce?
Per-operator learning curves, trajectory classifications (accelerating, plateaued, reverting), cohort movement analysis.
What enterprise question does it answer?
Are our operators getting better — and at what rate?
Turn improvement into measurable experiments.
What is it?
Instead of rolling out AI initiatives broadly and hoping for improvement, enterprises test targeted changes against specific operator benchmarks. Baseline → intervention → re-measure → compare → retain or discard.
What data is required?
Pre-intervention baseline + post-intervention re-measurement + declared target metric and follow-up window.
What does it produce?
Experiment records with target and non-target deltas, retention/discard recommendations, accumulated intervention evidence.
What enterprise question does it answer?
Did this intervention actually improve operator performance?
Evaluation should change what happens next.
What is it?
Every intervention declares a target metric, expected direction, follow-up window, and non-target metrics to monitor. The eval is repeated to determine whether the intervention improved measurable operator performance.
What data is required?
Pre-intervention eval + post-intervention eval + intervention record with declared target.
What does it produce?
Before/after comparison, target metric delta, non-target metric deltas, intervention effectiveness assessment.
What enterprise question does it answer?
What changed after we intervened?
Evidence labels enforced in code.
What is it?
Decision-use labels and evidence semantics enforced in the system. Every output carries a label declaring its epistemic status.
What data is required?
No additional data — governance metadata is embedded in every measurement, diagnosis, and intervention record.
What does it produce?
PROVEN · MEASURED · OBSERVED · DERIVED · HYPOTHESIS · PILOT · VALIDATION REQUIRED labels. DEVELOPMENTAL gates that route workflows, not people. ASSOCIATION labels on outcome joins — never CAUSATION.
What enterprise question does it answer?
What can we safely do with these results?