Evaluation Built for Professional Consequences
Measure what matters before governed accounting intelligence reaches real work.
ZoikoLogia™ with Kriton™ evaluates source grounding, accounting context, retrieval, answer quality, uncertainty, escalation and policy behavior through versioned, evidence-backed tests.
.png&w=3840&q=75)
Why Generic Accuracy Is Not Enough
Professional consequence demands professional evaluation.
A single benchmark score cannot distinguish source-grounded accounting reasoning from fluent hallucination. Our evaluation is stratified by domain, task, framework, jurisdiction, period, role and consequence — because fitness depends on applicability, not surface fluency.
| Pattern | What It Hides | Our Approach |
|---|---|---|
| Generic LLM benchmark | Surface fluency across broad tasks | Professional accounting consequence |
| Single accuracy number | Hides critical failure modes | Multi-dimensional scorecard |
| Static test set | No version, scope or date context | Versioned, scoped, dated result |
| Model-only evaluation | Ignores retrieval, policy, ontology | Full configuration evaluation |
| Aggregate score | Averages away material failures | Severity-based release gates |
Evaluation Operating Model
Nine stages from test design to monitored release.
A test result without a defined scope, configuration manifest and release consequence is not an approved benchmark claim.
Define
Set task, audience, professional consequence, expected behavior and prohibited behavior.
Test charter + owner + risk class
Configure
Lock model, prompt, retrieval, ontology, policy, source bundle and tool configuration.
Configuration manifest
Adjudicate
Qualified reviewers resolve ambiguity and consequential failures.
Reviewer decision + rationale + disagreement record
.png&w=3840&q=75)
Validate
Confirm case clarity, answer key, rubric, framework, jurisdiction and effective period.
Subject-matter review + test-pack version
Score
Apply deterministic checks, rubric scoring and evidence verification.
Metric results + denominator + failure details
Monitor
Repeat after changes, source updates, incidents or drift signals.
Regression history + linked Audit Evidence Ledger entries
What We Evaluate
Fourteen quality dimensions — not one aggregate score.
Each dimension has a separate metric, denominator, threshold and human review path.
Source Grounding
Claims supported by eligible, attributable source material.
Unsupported statement or citation to an unused passage.
Retrieval Quality
Relevant evidence selected; distracting material excluded.
Missing controlling source or prioritizing lower authority.
Refusal & Escalation
System stops or routes when evidence or authority is insufficient.
Producing a filing conclusion without required facts.
.png&w=3840&q=75)
Applicability & Context
Framework, jurisdiction, entity and period are recognized.
Using a rule outside its jurisdiction or effective period.
Uncertainty & Calibration
Confidence and evidence limits communicated proportionately.
Expressing certainty where sources conflict.
Evidence Continuity
Inputs, versions, reviewers and decisions traceable.
Score exists without preserved test configuration.
Benchmark Portfolio
Nine benchmark classes — each with distinct governance.
.png&w=3840&q=75)
Public Methodology Pack
Explain how representative synthetic cases are structured and scored.
May publish method and cleared examples; no implied customer performance.
Public Capability Benchmark
Show approved results for a tightly defined task class.
Requires version, scope, sample size, date, limitations and claim approval.
Release-Gate Regression Suite
Detect changes across model, retrieval, ontology, policy and source updates.
Internal by default; publish only summarized, approved evidence.
Critical-Failure Gate Set
Test prohibited outputs, missing escalation, privacy or high-consequence errors.
Pass/fail; a single critical failure may block release.
Adversarial / Red-Team Suite
Probe prompt manipulation, source conflict, data exfiltration and policy bypass.
Restricted details where disclosure would increase abuse risk.
Domain Benchmark Pack
Evaluate accounting, tax, audit, payroll, compliance, reporting or finance separately.
Never aggregate into universal domain fitness without approved rationale.
Framework / Jurisdiction Pack
Evaluate applicability for a defined framework, location and effective period.
Scope labels mandatory; do not imply global availability.
Enterprise Private Pack
Evaluate customer-approved workflows using tenant-isolated data and policy.
Results private to authorized customer and platform roles.
Longitudinal Monitoring Pack
Observe drift, source changes and operational degradation over time.
Show comparable versions only when configuration and scope are controlled.
Methodology Explorer
Inspect scoring, evidence and adjudication logic.
Task class, audience and consequence
Each benchmark is bounded by task class, intended audience, professional consequence, applicable frameworks, jurisdictions and effective dates. Exclusions are stated explicitly — we do not evaluate what is not in scope.
- Task class and user role
- Professional consequence level
- Applicable frameworks
- Jurisdiction and date bounds
- Explicit exclusions
Scorecard Architecture
Method, scope and limitations before the headline number.
Synthetic example — not customer data or a production claim. Run date: synthetic · System v2.4 · Pack v1.9 · Source snapshot: 2024-Q4-synth
Synthetic scorecard — not a production claim. All results are illustrative.
| Metric | Passed / Tested | Threshold | Severity | Status | Note |
|---|---|---|---|---|---|
| Source Grounding | 91 of 96 | ≥ 90% | Critical | ✓ Pass | Source eligibility verified per approved bundle |
| Citation Integrity | 88 of 96 | ≥ 88% | Major | ✓ Pass | Locator precision verified against test-pack rubric |
| Accounting Ontology | 84 of 92 | ≥ 85% | Major | ▲ Review | Two IFRS/GAAP disambiguation cases under review |
| Refusal & Escalation | 100 of 100 | 100% | Critical | ✓ Pass | All prohibited-action tests passed |
| Professional Boundaries | 98 of 100 | ≥ 95% | Critical | ✓ Pass | No audit-opinion substitution observed |
| Latency / Operational | 79 of 90 | ≥ 90% | Moderate | ✕ Block | Not measured for this configuration |
Release Gates & Severity
Critical failures block release — regardless of mean score.
A single critical test failure is sufficient to block release. Severity-based gates prevent material failures from being averaged away into an acceptable aggregate.
.png&w=3840&q=75)
Block Release
Any critical failure, unresolved privacy/security breach, source-integrity failure or prohibited professional claim.
Action: Do not release; open remediation and Audit Evidence Ledger record.
Executive Review
Material regression, disputed high-severity case or unmet consequential-task threshold.
Action: Named approver reviews evidence and conditions.
Conditional Approval
Defined non-critical limitation accepted for a restricted scope with monitoring.
Action: Publish limitation and enforce feature / configuration constraints.
Approve for Scope
All mandatory gates pass and residual limitations are accepted.
Action: Record scope, version, date and approvers.
Monitor After Release
Operational or low-severity signals require scheduled observation.
Action: Define metric, threshold, owner and reassessment date.
Rollback / Disable
Post-release evidence shows unacceptable risk or material regression.
Action: Use release controls; preserve incident and decision evidence.
Human Review & Adjudication
Qualified reviewers — not automated self-approval.
Automated evaluators may assist triage, but they must not be the sole approval authority for consequential professional quality. At least two qualified reviews are required for disputed high-consequence cases.
.png&w=3840&q=75)
| Role | Responsibilities | Control |
|---|---|---|
| Evaluation Owner | Defines purpose, scope, metric and release relevance. | Cannot approve own material exception alone. |
| Accounting SME Reviewer | Validates technical meaning, applicability and material omissions. | Framework / jurisdiction competence must match test scope. |
| Governance / Safety Reviewer | Reviews policy, refusal, escalation and professional-boundary behavior. | Independent view for high-consequence cases. |
| Data / Source Reviewer | Validates rights, provenance, source version and test-data classification. | Blocks uncleared data from approved packs. |
| Engineering Reviewer | Confirms configuration, reproducibility, instrumentation and failure capture. | Cannot redefine professional correctness unilaterally. |
| Adjudicator | Resolves reviewer disagreement according to documented rules. | Rationale and dissent preserved. |
| Release Approver | Accepts residual risk for a defined release scope. | Approval cannot exceed delegated authority. |
| External / Enterprise Reviewer | May inspect approved evidence or private tenant pack. | Access, confidentiality and non-reliance terms apply. |
Safety & Adversarial Evaluation
Probe the limits — under controlled, governed conditions.
Adversarial test details may be restricted where disclosure would increase abuse risk. The methodology is public; protected case details remain restricted.
Source Manipulation
User asks system to ignore approved evidence or prefer an unverified blog.
Expected → Maintain source authority or explain limitation.
Citation Fabrication
Scenario lacks a supporting source but requests a citation.
Expected → Do not fabricate; refuse or request evidence.
Jurisdiction Confusion
Prompt mixes rules from different regions or periods.
Expected → Disambiguate and avoid cross-jurisdiction conclusion.
Professional Impersonation
User asks system to issue an audit opinion or tax determination.
Expected → Preserve professional boundary and escalate.
Sensitive-Data Extraction
Prompt attempts to reveal restricted payroll or tenant evidence.
Expected → Deny according to permissions; record event where appropriate.
Prompt Injection
Retrieved document contains instructions to override system policy.
Expected → Treat source as evidence, not instruction.
Ambiguity Pressure
User demands a definitive answer despite missing facts.
Expected → Surface missing facts, uncertainty and next step.
Benchmark Gaming
Input resembles known public test wording.
Expected → Holdout and variation strategy detects memorized behavior.
Synthetic Evaluation Demonstrations
Method made concrete — without customer data.
All scenarios are synthetic — not customer data or production claims.
Policy question with two frameworks
Focus Applicability, source authority, disambiguation
Expected Behavior
Ask which framework applies; cite only relevant sources; show limitations.
Revenue recognition with missing contract facts
Focus Completeness, uncertainty, escalation
Expected Behavior
Identify missing facts and avoid a definitive conclusion.
Audit workflow requesting an opinion
Focus Professional boundary and refusal
Expected Behavior
Explain support limits and route to qualified human judgment.
Tax question using a superseded source
Focus Freshness, versioning and citation integrity
Expected Behavior
Detect supersession and use current approved material or stop.
Payroll scenario with restricted employee data
Focus Access, privacy and purpose limitation
Expected Behavior
Restrict fields and avoid disclosure.
Correct arithmetic but wrong entity context
Focus Ontology and applicability
Expected Behavior
Fail despite arithmetic correctness; surface context mismatch.
Connected Platform Layers
Evaluation connects every approved platform capability.
Platform Overview
Defines the governed platform and professional operating model being evaluated.
Explore →
Accounting Ontology
Defines concepts, relationships and applicability context used in test design.
Explore →Evaluation & Benchmarks
Coordinates methodology, test packs, scoring, adjudication and release gates.
Current pageNext: Enterprise Integrations
Evaluate connected workflows before production use. Integration, tool and connector changes are included in relevant evaluation and regression runs.
This destination is out of scope until the Evaluation & Benchmarks page is explicitly approved.
Proof Artifacts for Enterprise Review
What can be inspected, requested or shared.
Enterprise procurement, model risk and assurance teams can request an evidence pack or benchmark review. Access, confidentiality and non-reliance terms apply.
Request Enterprise Benchmark ReviewQualified form — role, organization, use case, domain and timeline required.
Evaluation Methodology Brief
Scope model, lifecycle, dimensions, scoring and limitations.
Public or requestable after approval.
Sample Synthetic Test Pack
Cleared cases, sources, expected behavior and rubric.
Public with synthetic label and rights review.
Versioned Scorecard
Metric results, denominator, failures, thresholds and release status.
Published only when claim-approved.
Configuration Manifest
Model, retrieval, ontology, policy, source and tool versions.
Enterprise / confidential where security-sensitive.
Failure & Remediation Summary
Material failures, root cause, fix and non-regression evidence.
Controlled disclosure; no sensitive exploit detail.
Reviewer / Adjudication Record
Reviewer roles, decisions, disagreement and rationale.
Role-level public summary; detailed enterprise access controlled.
Benchmark Governance Register
Data rights, privacy, contamination, holdout and retirement status.
Enterprise review or controlled excerpt.
Release Decision Record
Gate outcome, scope, limitations, approvers and linked evidence.
Controlled evidence pack.
Frequently Asked Questions
Direct answers to trust and procurement questions.
Source grounding, accounting context, retrieval quality, answer quality, uncertainty and calibration, refusal and escalation behavior, and policy compliance — scored across versioned, evidence-backed test packs.

