Measure AI behaviour against versioned cases.

DecarbDesk uses structured evaluation to test outputs, compare configurations, and detect regressions. Deterministic workflow controls decide whether an output continues, pauses, or requires review.

Treat prompts and model settings as configurations to test.

AI behaviour can change when the model, input format, source data, or prompt configuration changes. Manual spot checks may miss a regression or apply different standards from one reviewer to another.

DSPy provides a framework for defining language-model programs, evaluating them against examples, and optimizing configurations against a stated objective. We use those capabilities inside a broader workflow that also includes access rules, deterministic routing, human approvals, and audit records.

DSPy does not replace the control layer. It helps measure and improve model-dependent behaviour; the surrounding application enforces what happens next.

System boundary

Three layers with different responsibilities.

01

Orchestration

Runs triggers, routing, permissions, approvals, limits, logging, and stop conditions. This layer follows explicit rules.

02

Evaluation

Scores model-dependent outputs against defined metrics and reference cases. DSPy supports evaluation and configuration optimization here.

03

Model operation

Extracts, classifies, drafts, summarizes, or recommends. Its output remains subject to the first two layers.

Record what was tested and what happened next.

A score has limited value without the version, threshold, and resulting decision.

Reference set
Versioned examples that represent normal cases, edge cases, and known failures.
Metric
A defined test for accuracy, format, policy compliance, consistency, or another agreed objective.
Threshold
The minimum score or rule result required to continue.
Decision
Pass, route for review, pause the workflow, or block a deployment.
Trace
Input version, model configuration, output, score, reviewer decision, and timestamp.

Checks tied to an operational response.

Document extraction

Test: Compare extracted totals with line items and required fields.

Response: Route mismatches or missing fields to review before data is written to the target system.

Intake classification

Test: Score the proposed category against labeled examples and policy rules.

Response: Send low-confidence or high-impact cases to manual triage.

Draft correspondence

Test: Check source accuracy, required language, prohibited content, and tone criteria.

Response: Hold the draft for correction or human approval when a check fails.

Track quality with throughput and exceptions.

Operations reporting should show score movement, failed checks, review volume, and configuration changes alongside ordinary workflow performance.

Hours returned

to your team each week

Collection cycle

improvement over baseline

Manual errors

reduced by automation

Forecast window

accuracy vs. actuals