Baseline. Benchmark. Intervene. Re-evaluate.
25–100 users. 30 days. Operator telemetry, cohort analysis, bespoke eval design, internal benchmarking, external reference comparisons, workflow-level questions, targeted intervention recommendations, and a re-evaluation plan.
Your people. Your workflows. Your evals. Your benchmarks. No productivity scores. No employee termination use.
Run a 30-Day Operator Eval See What We BenchmarkSeven steps. 30 days.
- Instrument: Connect operator telemetry. Configure privacy boundaries.
- Baseline: Establish operator performance baseline across the cohort.
- Bespoke Evals: Build company-specific evals around your workflows, roles, and models.
- Benchmark: Compare operators, teams, and workflows against internal and external benchmarks.
- Diagnose: Identify capability concentration, workflow friction, model/operator fit, dependency risk.
- Intervene: Design targeted interventions — training, workflow redesign, model placement, governance controls.
- Re-evaluate: Re-measure to determine what changed. Compare target and non-target deltas.
What you get.
Per-operator performance profile, field position, percentile band, archetype, volatility.
Who your strongest operators are — structurally, not by volume.
Median, top 25%, top 10%, top 5%, top 1%. Where each operator sits.
How performance spreads across your workforce. Shape, skew, gaps, clusters.
Where operators, models, and tools fit specific workflow stages — and where they don't.
Which operators operate AI the same way. Archetypes, peer groups, hidden capability.
How capability is distributed across teams using the same AI stack.
Whether advanced practice is organizational or isolated in a few people.
Which models amplify which operators. When to change the model vs develop the operator.
How fast operators are improving. Acceleration, plateau, reversion, emergence.
Targeted recommendations tied to measured gaps. Re-evaluation plan with declared target metrics and follow-up windows.
Six benchmark types. One system.
Common metrics usable across organizations. Leverage, yield, construction, field position.
Compare operators and teams within the organization. Percentile bands, team comparisons, department comparisons.
Compare against relevant reference populations. SigRank operator field. Field position, distance from field bands.
Built around organization-specific workflows, roles, and tasks. Your work, your benchmarks.
Compare performance over time. Acceleration, plateau, reversion, cohort movement.
Compare pre- and post-intervention performance. Target and non-target deltas. Retain or discard.
12 pilots. One evaluation system.
Each pilot answers a different enterprise question through operator evals and performative benchmarks. Choose by outcome or build your own from 15 evaluation families.
| # | Pilot | Question | Evals | Level |
|---|---|---|---|---|
| 1 | AI Workforce Operating Baseline | What does our AI workforce actually look like? | 5 | 1 |
| 2 | AI Capability Distribution | Do we have organizational capability or a few power users? | 2 | 1 |
| 3 | AI Adoption & Adaptation | Are people adapting, or simply using the tools? | 2 | 1 |
| 4 | AI Training Evaluation | Did our training actually change how people operate AI? | 2 | 2 |
| 5 | Model / Tool Evaluation | What changes when we introduce Model A, Model B or Tool X? | 2 | 2 |
| 6 | Agent Adoption | Is agent capability becoming organizational or staying concentrated? | 2 | 2 |
| 7 | Workflow Diagnostic | Is the constraint the operator, the tool or the workflow? | 2 | 1 |
| 8 | Team AI Operating Comparison | Why do teams using the same AI stack operate differently? | 3 | 1 |
| 9 | Experiments | What changed when we introduced X? | 2 | 2 |
| 10 | Monitor | How is our AI operating population changing month over month? | 3 | 1 |
| 11 | Meta-Pilot | Does this metric predict something we already care about? | 2 | 1 · ASSOCIATION |
| 12 | Vendor / Consultancy Verification | Are our outsourced AI vendors actually delivering quality? | 3 | 2 |
Choose by outcome. Or build your own.
The pilot configurator lets you answer "what outcomes are you looking for?" and get a recommended pilot — or select individual evaluation families à la carte.
Pick a commercial pilot from the catalog. Each comes pre-packaged with the right eval families, deployment level, and governance metadata.
- 12 pre-packaged pilots
- Automatic eval family selection
- Deployment level matched to data requirements
- Governance metadata embedded
Select from 15 evaluation families à la carte. The configurator validates compatibility and warns on unimplemented evals.
- 15 eval families (13 implemented)
- 10 compatibility validation rules
- Configurable gates, cohort, workflow
- Save / load configurations as JSON
Available via CLI, TUI, and MCP. Every configuration carries governance metadata: synthetic flag, decision-use labels, ASSOCIATION semantics for outcome joins, HYPOTHESIS semantics for diagnoses.
15 eval families. 13 implemented. 2 in development.
Each pilot is composed of evaluation families. You can choose a pre-packaged pilot or build your own à la carte. The configurator validates compatibility across 10 rules.
| ID | Name | Status |
|---|---|---|
| EVAL-001 | Operator Baseline | full |
| EVAL-002 | Usage vs Operation Divergence | full |
| EVAL-003 | Context Architecture | partial |
| EVAL-004 | Longitudinal Movement | partial |
| EVAL-005 | Platform / Model Sensitivity | full |
| EVAL-006 | Cohort Composition | full |
| EVAL-007 | Intervention Response | full |
| EVAL-008 | Workflow Stage Fit | full |
| EVAL-009 | Team Composition | partial |
| EVAL-010 | Capability Dependency Risk | partial |
| EVAL-011 | Development Engine | full |
| EVAL-012 | Experiment as Product | full |
| EVAL-013 | Org AI Topology | not implemented |
| EVAL-014 | Operator Similarity Search | not implemented |
| EVAL-015 | AI Learning Curve | partial |
Three levels. Matched to data requirements.
Best for Baseline, Capability Distribution, basic Adaptation, Team Comparison and Monitor. Uses the current measurement core and does not require prompt content.
Best for Training Evaluation, Model/Tool Evaluation, Agent Adoption, richer Workflow Diagnostics and Experiments. Adds timestamps, sessions, model/tool/agent data, acceptance, retries, errors and richer provider fields where available.
Best for continuous internal measurement, customer-defined thresholds, authoritative lineage, governed state and later MO§E§ deployment.
Six packages. From baseline to production.
Map your AI workforce.
Find where capability is concentrated, stalled or structurally constrained.
Measure what changes when you introduce a model, tool, agent, workflow or training program.
Track how your AI operating population evolves.
Validate that the metric predicts something you care about. ASSOCIATION — never CAUSATION.
Govern and own the measurement infrastructure once the pilot becomes production.
Three surfaces. One service.
The pilot does not require a full enterprise dashboard. Delivery may include CLI, TUI, MCP-style interface, structured reports, bespoke analysis, and executive summary.
The reproducible operational layer. Ingest, score, cohort, diagnose, workflow, intervene, verify, export, configure.
Rich console for live review. Pilot overview, cohort summary, operator profile, distributions, divergence, diagnostic queue, workflow map, interventions, before/after, export, configuration.
Read-first analytical tools for AI assistants. 16 tools including pilot configuration, validation, and reporting. Governance annotations on every response.
A dashboard comes later, only when repeated pilots reveal stable recurring views.
What the pilot is not.
MO§ES™ does not claim that a higher metric score means a better employee, higher job performance, greater productivity, or better business outcomes without separate validation.
The system works from telemetry and structural signals, not prompt content. Cohort-level reporting by default. No adverse employment action in pilot.
We measure operator change directly. We test business impact separately. A correlation between an intervention and a business metric is not proof that the intervention caused the business change.