Environments to teach models
real back-office work

Frontier agents still fail real back-office work. Our environments make them better at it.

Why today's benchmarks miss the back office

Agent performance is usually measured in agent-friendly settings: broad system access, wide permissions, and clean tooling built to be called programmatically. The conditions in a European back office differ on each of those points. Permissions are issued per role and reviewed, so an account reaches the cases it is assigned and little beyond them. Core systems were procured before integration was a requirement and are reached through their own screens, not through an interface built for calling them. Actions are logged because the organisation has to be able to demonstrate who changed what, which also means a wrong action is not undone by repeating it. Regulation and accumulated technical debt sustain this arrangement jointly.

Cost vs average partial score scores from our data sample, available on request 0% 25% 50% 75% 100% $0 $5 $10 $15 $20 $25 cost per attempt mean partial score gemini-3.1-pro gemini-3.5-flash gpt-5.5 gpt-5.6-sol opus-5 opus-4.8 fable-5

Enterprises will not rebuild that estate to make it easier for an agent to work in. In practice a model reaches the right answer and then fails to commit it, because committing it takes a long sequence of small operations, each able to fail on its own: carrying a figure between systems, separating near-identical records, obeying an unexplained order of entry, and remembering what has already been done when no system keeps count.

What our environments offer
01

Workplace realism

A benchmark gives a shell, one application in a clean profile and root access. The work runs on a locked-down virtual desktop behind SSO and a firewall, across an office bundle and a vendor suite.

applications open at once 1 avg. benchmark 6 our environments
02

Moving state

Benchmarks often hand over one prompt and collect one artefact. A working day changes the state underneath the task, and an agent that passes at the end can still fail the day.

state changes in one day 09:3011:15 13:4015:50 17:00
03

Domain coverage

Most environments are built for computer and mathematical work. We work strictly on office and administrative support.

tasks per million workers 1,658 computing 175 office admin

Those operations are not the only place a run goes wrong. An agent may reach the correct total by a route the rules forbid, or use the most accessible document instead of the one that governs. An environment that takes the operational burden away tests the calculation and not the task.

A core system in a regulated industry is specified in a tender, contracted for years, and then migrated as rarely as possible, because a migration puts the operational record at risk. The software a clerk opens tomorrow was chosen before the current generation of models existed and will still be in place for the generations that follow. Labs will have to ship models that act in the same workplaces people have to work in, through the same systems.

Operating-cost settlement, worked through the screens GPT-5.6 Sol, xhigh · 50% score on this task
Computer use over a 17-tool MCP server: screenshot, click, type, fill_field, navigate. One login carries across all six systems, which the agent reaches only through their own screens.
Figure 1: One recorded run, every step. GPT-5.6 Sol reached a score of 0.5 and lost the rest on judgment rather than mechanics: it booked an unrecorded invoice to the wrong account, and read a missing quarterly instalment as present.

The capabilities our environments teach

Frontier agents get a lot of back-office work right, the domain reasoning included, and still lose the case. From our data sample we identified five capabilities a case demands between the evidence and the close, and our environments teach them. Our verifiers grade each one on its own.

One case, from the evidence to the close reads Source authority applies Rule application decides Authorisation records Evidence references confirms Confirmed outcome filled: share of recorded runs that got the stage right, one dot per tenth Source authority The same information sits in several places: a scanned original, a machine-read extract of that scan, a provider's portal. The domain rules define which source governs. Rule application The governing rule can be a regulation or a company procedure. Finding it and committing the value it produces are separate steps, and runs name the rule and then book another figure. Authorisation Every action runs on the rights the role holds. A booked addition the policy does not cover voids the whole settlement, and a case the rules reserve for escalation has to leave the desk. Evidence references Every decision carries a rule reference, an evidence reference and a note. A decision that is right on substance and thin on its basis fails the same as a wrong one. Confirmed outcome A form can accept an entry and change nothing, and the screen confirms the request rather than the state. The agent reads back what the system now holds.
Figure 2: The five capabilities a case demands, in the order the work happens. The dots are the share of recorded runs that lost nothing in that capability, counted from the graded dimensions of every trial.

Realistic is what makes them hard

We build environments that reproduce a workplace, and the difficulty comes from the reproduction rather than from anything we add. Most of the build goes into understanding one workflow: what a clerk, a manager or an analyst actually does, in what order, under whose authority.

Environments are mirrors of existing workplaces, with the applications rebuilt screen by screen where that is what fidelity requires. For the German statutory health insurers that means the case-management suite the sector runs on, reproduced with the process it carries, so what the agent has to get right is the process those offices actually follow.

Our quality promise

Every task is validated by domain experts, from the technical build through the process it reproduces to the domain knowledge it rests on. An environment moves through three stages, and each ends in a gate that can send it back.

The three stages
Stage What happens What it has to prove
First review Vertical and role selection, the workflow mapped end to end, the workplace scaffold chosen, cases generated against the real regulation. The role and the process boundaries survive review against independent evidence and the domain expert's reading of them.
Second review The environment is implemented, the reward is designed against committed state, and the task is packaged as a runnable benchmark task. A reference solution scores full reward through the same interface an agent uses, and doing nothing scores zero.
Merge Frontier models from different families are run at production settings and the traces are read step by step. The environment separates models rather than failing all of them for one avoidable reason. Anything a trace shows to be unfair is fixed and the trials are rerun.
Table 1: The second gate is the one that catches most problems: an environment a reference solution cannot solve is broken, and one a no-op can solve is not measuring anything.

Why ellamind and not a global data vendor

We are based in Germany, with a broad network and direct access to practitioners in the industries we serve: German public authorities, insurance companies, retailers and others that carry a large back office. We build agent systems for them, and we build environments to measure performance before deployment, to see where the risk sits and to improve the agents before they go into business processes that carry real consequences.

What makes ellamind a trustworthy partner

We are involved in some of the largest German and European foundation-model projects, among them OpenEuroLLM and SOOFI, where we provide evaluations from early pretraining through to post-training.

We publish a large part of that work. propella-1 annotates pretraining data at scale for data curation. We release data mixes, annotated pretraining data and benchmark suites, among them a German base-model evaluation suite. We are an official data partner for Terminal-Bench 3 and part of our research team helped out with QA.

If you are interested in a sample package:

Get in touch