ellaverse

Agentic evaluation environments

Real-world environments for enterprise AI agents.

Specialized, domain-rich environments that mirror actual knowledge work, not generic coding tasks. Each environment recreates a real business workflow, complete with domain-specific GUIs, authentic documents, regulatory rules, and ground-truth labels.

ellaverse PKV claims environment: inbox with 27 open claims sorted by SLA priority

Benchmarks for the work that actually matters.

Unlike generic coding benchmarks, ellaverse targets knowledge work in regulated industries: insurance claims, product data compliance, legal review.

Domain-specific

Built on real business processes, regulations, and domain rules. Not synthetic tasks.

Production-realistic

Full GUIs, real documents, databases, and decision workflows. Not toy setups.

Rigorously validated

Ground-truth labels, scoring rubrics, and reproducible baselines for every environment.

The environment catalog.

Every environment ships with verified ground-truth answers, so accuracy is measured objectively and progress is trackable over time.

Live

PKV Claims Processing

Healthcare · Germany

AI agents review German private health insurance claims against GOÄ billing rules in a GUI-based claims processing workspace.

Insurance GOÄ Billing Document Processing GUI-based
27 cases
283 decisions
18 error types
Live

Retailer Product Listing QA

Retail · Global

AI agents act as key account managers on a retail partner portal, reviewing flagged product listings across four categories against category schemas, EU product safety rules, and vendor data.

E-Commerce Product Data Quality Multi-App Email
33 cases
33 decisions
7 error types
Coming soon

Commercial Underwriting

Insurance · Europe

AI agents evaluate SME insurance applications, assess risk profiles, and produce binding and coverage recommendations.

Underwriting Risk Assessment SME Insurance
Coming soon

KFZ Liability Claims

Insurance · Germany

AI agents process motor vehicle liability claims, assess fault allocation, and calculate damage settlements against German insurance regulations.

Motor Insurance Liability Assessment Damage Estimation
Coming soon

Legal Document Review

Legal · Global

AI agents review commercial contracts for clause compliance, identify risk provisions, and produce structured redline summaries.

Contract Review Compliance Risk Analysis

Inside the live environments.

PKV Claims Processing

Germany's private health insurance (PKV) reimburses patients directly, so every invoice is checked line by line against the GOÄ fee schedule: 1,634 procedure codes with multiplier thresholds and exclusion rules. This environment recreates that daily workflow with 27 cases and 283 invoice lines against a known ground truth, including deliberate extraction errors an agent has to catch.

What this tests

  • Navigate a domain-specific GUI: inbox, case details, embedded PDF viewer, and decision forms.
  • Apply complex billing rules: multiplier thresholds (Steigerungsfaktor), mutually exclusive codes (Ausschlussziffern), and justification requirements.
  • Cross-reference sources: the structured data contains deliberate extraction errors in 18 of 27 cases, and the PDF invoice is the ground truth.
  • Decide per invoice line: approve, reduce to a specific amount, or reject, with the correct monetary values.
PKV claims case detail: PDF invoice alongside structured data with billing rule flags

Retailer Product Listing QA

Agents act as key account managers on a retail partner portal: 33 flagged product listings across four categories have to be triaged against category schemas, EU product safety rules (GPSR), and vendor data, while emails from the team lead and vendors change priorities mid-session.

What this tests

  • Operate a multi-application workspace: partner portal plus email inbox, with new information arriving mid-task.
  • Apply category-specific schemas with required fields, data types, and allowed values.
  • Triage under ambiguity: correct the listing, query the vendor, escalate, or deactivate.
  • Extract structured data from free-text descriptions and product images, across 7 distinct error types.
Retail partner portal: flagged product listings with data quality warnings

How frontier models score today.

Benchmark runs from the ellaverse evaluation harness. Even frontier models leave plenty of headroom on realistic enterprise work: the best model resolves less than half of the PKV claims environment.

PKV Claims Processing · best score per model
Model Agent harness Score
gpt-5.4 codex 48.1 %
claude-sonnet-4-6 claude-code 37.0 %
gpt-5.3-codex codex 37.0 %
claude-opus-4-6 claude-code 29.6 %
gemini-3.1-pro gemini-cli 24.8 %
gemini-3-flash gemini-cli 14.8 %

On Retailer Product Listing QA, claude-opus-4-6 reaches 85.2 % on the medium-difficulty task; the easy variant is solved at 100 % by claude-opus-4-6 and gemini-3.1-pro.

Snapshot from ellaverse benchmark runs, March 2026. Scoring: correct tasks / total tasks (PKV claims); weighted triage, attribute accuracy, and coverage (retailer listing).

Test your agents before production does.

We'll walk you through the environments on your use case, from the first benchmark run to a customized environment for your own processes.