Claude Code vs Codex: Codex Plans, Claude Debugs
I had a rule: Claude plans, Codex debugs. So I replayed four real engineering tasks against both agents, three times each. 24 runs and 468 ratings later, the rule was backwards.
Insights on AI evaluation, experimentation, and building reliable AI products.
I had a rule: Claude plans, Codex debugs. So I replayed four real engineering tasks against both agents, three times each. 24 runs and 468 ratings later, the rule was backwards.
Azure is retiring GPT-4o, and our customer's production response times had already degraded to the point of raising timeouts. Here is how we swapped the model in 48 hours, validated, red-teamed, and approved, using elluminate.
A German pun most Germans nail in a second, yet frontier LLMs botch it. GRIPS is a broad, native German benchmark spanning reasoning, idioms, puzzles, and wordplay, with tasks built to require real reasoning and native-level German capabilities to solve.
A health insurer's vision model read 94% of scanned member documents correctly. For numbers that set what people pay each month, that isn't enough. Here's how an evaluation loop in elluminate found the six failures and fixed them, one experiment at a time, all the way to 107/107.
EU AI Act technical documentation under Annex IV is mandatory, has no template, and usually lands on one overworked person's desk. We're shipping elladoc, elluminate's compliance documentation assistant, in beta: a guided Classify → Document → Evaluate → Report workflow that shows the legal basis for every field and ties the claims to real evaluation evidence.
The EU AI Act has a clause that turns teams who only wrote a system prompt into the provider of a high-risk AI system — no shipping required. How Article 25(1)(c) works, what Annex IV then asks of you, and why settling your role first makes the rest of the documentation fall into place.
We turned the internet game 'Explain a Film Plot Badly' into an LLM evaluation across five models. Telling a model it was a world-class movie expert made it worse every single time — and so did upgrading to a newer model. Here's why, and why you only catch it with an eval.
What a Forward Deployed Engineer actually does at ellamind: onboarding customers onto elluminate, carrying the friction they don't file back into the roadmap, and closing the loop between a customer's reality and the product.
A CLAUDE.md file's value isn't the facts it supplies. It's the workflow it imposes. We ran Claude Opus 4.6 on ten Django tasks with and without our project's instructions, and the cost gap came almost entirely from one workflow choice the file encodes.
When a German health care provider needed to replace the LLM powering its customer-facing AI platform, structured evaluation turned a risky model switch into a data-driven decision. Uncovering fabricated personal data, leaked system instructions, and costly wrong reimbursement decisions along the way.
How N+1 evaluation catches the failures that single-turn testing misses, and how to set it up in elluminate.
We re-ran OpenAI's FrontierScience benchmark on GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro using elluminate. Here's where the latest frontier models stand on the hardest public science benchmark.
How structured evaluation turned a capable but unreliable AI agent into one that processes PKV claims at 93% accuracy. Same model, same cases, different instructions.
Quality ownership drifts when AI systems fail quietly. Learn why plausible-sounding outputs are dangerous, what ownership actually means for AI teams, and how to scale evaluation rigor with risk.
Running Kubernetes in production on multiple cloud providers means juggling OpenTofu configurations, Helm charts, and deployment pipelines. Here's how we use Claude Code as an infrastructure copilot with safety guardrails, custom skills, and encoded domain knowledge.
We tested 168 sensitive China-related topics across 10 LLMs. One Chinese model matched GPT-5.2 and Claude. Another rewrote the Tiananmen massacre as state-approved fiction.
A complete framework for RAG evaluation covering test set design, targeted criteria for retrieval and generation, experiment analysis, and continuous production monitoring.
Your evaluations say your AI is perfect. You know it's not. Here's how we used MCP to iterate rapidly and surface real limitations.
Import your Langfuse datasets directly into elluminate. Turn production traces into structured evaluations - no export scripts or CSV wrangling required.
Binary pass/fail evaluations beat Likert scales for LLM and agent evaluation. Here's why, and how to keep nuance without the inconsistency.
How we built a test set for a German health insurer's AI search—from 50 real user queries to 80 cases, 57 experiments, and a pass rate that climbed from 35% to over 80%.
Learn how to systematically test and improve your AI prompts using elluminate's evaluation platform. Walk through a complete example using pizza toppings to understand prompt templates, collections, criteria, and experiments.
See how our products can help you evaluate, deploy, and monitor AI agents with confidence.