Environments to teach models
real back-office work
Frontier agents still fail real back-office work. Our environments make them better at it.
Why today's benchmarks miss the back office
Agent performance is usually measured in agent-friendly settings: broad system access, wide permissions, and clean tooling built to be called programmatically. The conditions in a European back office differ on each of those points. Permissions are issued per role and reviewed, so an account reaches the cases it is assigned and little beyond them. Core systems were procured before integration was a requirement and are reached through their own screens, not through an interface built for calling them. Actions are logged because the organisation has to be able to demonstrate who changed what, which also means a wrong action is not undone by repeating it. Regulation and accumulated technical debt sustain this arrangement jointly.
Enterprises will not rebuild that estate to make it easier for an agent to work in. In practice a model reaches the right answer and then fails to commit it, because committing it takes a long sequence of small operations, each able to fail on its own: carrying a figure between systems, separating near-identical records, obeying an unexplained order of entry, and remembering what has already been done when no system keeps count.
Workplace realism
A benchmark gives a shell, one application in a clean profile and root access. The work runs on a locked-down virtual desktop behind SSO and a firewall, across an office bundle and a vendor suite.
Moving state
Benchmarks often hand over one prompt and collect one artefact. A working day changes the state underneath the task, and an agent that passes at the end can still fail the day.
Domain coverage
Most environments are built for computer and mathematical work. We work strictly on office and administrative support.
Those operations are not the only place a run goes wrong. An agent may reach the correct total by a route the rules forbid, or use the most accessible document instead of the one that governs. An environment that takes the operational burden away tests the calculation and not the task.
A core system in a regulated industry is specified in a tender, contracted for years, and then migrated as rarely as possible, because a migration puts the operational record at risk. The software a clerk opens tomorrow was chosen before the current generation of models existed and will still be in place for the generations that follow. Labs will have to ship models that act in the same workplaces people have to work in, through the same systems.
The capabilities our environments teach
Frontier agents get a lot of back-office work right, the domain reasoning included, and still lose the case. From our data sample we identified five capabilities a case demands between the evidence and the close, and our environments teach them. Our verifiers grade each one on its own.
Realistic is what makes them hard
We build environments that reproduce a workplace, and the difficulty comes from the reproduction rather than from anything we add. Most of the build goes into understanding one workflow: what a clerk, a manager or an analyst actually does, in what order, under whose authority.
Environments are mirrors of existing workplaces, with the applications rebuilt screen by screen where that is what fidelity requires. For the German statutory health insurers that means the case-management suite the sector runs on, reproduced with the process it carries, so what the agent has to get right is the process those offices actually follow.
Our quality promise
Every task is validated by domain experts, from the technical build through the process it reproduces to the domain knowledge it rests on. An environment moves through three stages, and each ends in a gate that can send it back.
| Stage | What happens | What it has to prove |
|---|---|---|
| First review | Vertical and role selection, the workflow mapped end to end, the workplace scaffold chosen, cases generated against the real regulation. | The role and the process boundaries survive review against independent evidence and the domain expert's reading of them. |
| Second review | The environment is implemented, the reward is designed against committed state, and the task is packaged as a runnable benchmark task. | A reference solution scores full reward through the same interface an agent uses, and doing nothing scores zero. |
| Merge | Frontier models from different families are run at production settings and the traces are read step by step. | The environment separates models rather than failing all of them for one avoidable reason. Anything a trace shows to be unfair is fixed and the trials are rerun. |
Why ellamind and not a global data vendor
We are based in Germany, with a broad network and direct access to practitioners in the industries we serve: German public authorities, insurance companies, retailers and others that carry a large back office. We build agent systems for them, and we build environments to measure performance before deployment, to see where the risk sits and to improve the agents before they go into business processes that carry real consequences.
What makes ellamind a trustworthy partner
We are involved in some of the largest German and European foundation-model projects, among them OpenEuroLLM and SOOFI, where we provide evaluations from early pretraining through to post-training.
We publish a large part of that work. propella-1 annotates pretraining data at scale for data curation. We release data mixes, annotated pretraining data and benchmark suites, among them a German base-model evaluation suite. We are an official data partner for Terminal-Bench 3 and part of our research team helped out with QA.
If you are interested in a sample package:
Get in touch