Domain-specific
Built on real business processes, regulations, and domain rules. Not synthetic tasks.
Agentic evaluation environments
Specialized, domain-rich environments that mirror actual knowledge work, not generic coding tasks. Each environment recreates a real business workflow, complete with domain-specific GUIs, authentic documents, regulatory rules, and ground-truth labels.
Unlike generic coding benchmarks, ellaverse targets knowledge work in regulated industries: insurance claims, product data compliance, legal review.
Built on real business processes, regulations, and domain rules. Not synthetic tasks.
Full GUIs, real documents, databases, and decision workflows. Not toy setups.
Ground-truth labels, scoring rubrics, and reproducible baselines for every environment.
Every environment ships with verified ground-truth answers, so accuracy is measured objectively and progress is trackable over time.
Healthcare · Germany
AI agents review German private health insurance claims against GOÄ billing rules in a GUI-based claims processing workspace.
Retail · Global
AI agents act as key account managers on a retail partner portal, reviewing flagged product listings across four categories against category schemas, EU product safety rules, and vendor data.
Insurance · Europe
AI agents evaluate SME insurance applications, assess risk profiles, and produce binding and coverage recommendations.
Insurance · Germany
AI agents process motor vehicle liability claims, assess fault allocation, and calculate damage settlements against German insurance regulations.
Legal · Global
AI agents review commercial contracts for clause compliance, identify risk provisions, and produce structured redline summaries.
Germany's private health insurance (PKV) reimburses patients directly, so every invoice is checked line by line against the GOÄ fee schedule: 1,634 procedure codes with multiplier thresholds and exclusion rules. This environment recreates that daily workflow with 27 cases and 283 invoice lines against a known ground truth, including deliberate extraction errors an agent has to catch.
What this tests
Agents act as key account managers on a retail partner portal: 33 flagged product listings across four categories have to be triaged against category schemas, EU product safety rules (GPSR), and vendor data, while emails from the team lead and vendors change priorities mid-session.
What this tests
Benchmark runs from the ellaverse evaluation harness. Even frontier models leave plenty of headroom on realistic enterprise work: the best model resolves less than half of the PKV claims environment.
| Model | Agent harness | Score |
|---|---|---|
| gpt-5.4 | codex | 48.1 % |
| claude-sonnet-4-6 | claude-code | 37.0 % |
| gpt-5.3-codex | codex | 37.0 % |
| claude-opus-4-6 | claude-code | 29.6 % |
| gemini-3.1-pro | gemini-cli | 24.8 % |
| gemini-3-flash | gemini-cli | 14.8 % |
On Retailer Product Listing QA, claude-opus-4-6 reaches 85.2 % on the medium-difficulty task; the easy variant is solved at 100 % by claude-opus-4-6 and gemini-3.1-pro.
Snapshot from ellaverse benchmark runs, March 2026. Scoring: correct tasks / total tasks (PKV claims); weighted triage, attribute accuracy, and coverage (retailer listing).
We'll walk you through the environments on your use case, from the first benchmark run to a customized environment for your own processes.