GRIPS: A Native-Level German Reasoning Benchmark for LLMs

Maximilian Idahl

Maximilian Idahl

10 min read

Prompt

Vervollständige dieses bekannte Wortspiel: „Egal wie neu du bist, Manuel ist ___.“

OpenAI

GPT-5.5

Neuer

Anthropic

Claude Opus 4.8

Neuer

Google

Gemini 3.5 Flash

Neuer

DeepSeek

DeepSeek V4 Pro

Neuer

Qwen

Qwen3.7 Max

Neuer

Kimi

Kimi K2.6

Neuer

NVIDIA

Nemotron 3 Ultra

manuell

Zhipu

GLM 5.2

manuell

MiniMax

MiniMax M3

dabei

Mistral

Mistral Medium 3.5

älter

That’s a well-known German pun, and the answer is Neuer, the goalkeeper Manuel Neuer, whose name sounds exactly like the comparative of neu (“newer”). Most Germans get it in a second. Six of the models above land on it too. The rest miss in telling ways: two settle for „manuell“ (right phonetic neighborhood, wrong target), and the others drift to „dabei“ or „älter“: fluent, confident, and wrong. Models rarely say “I don’t know”; they hand back a clean wrong answer instead. And this isn’t a one-off gotcha. Puns are just one category of German puzzles that trip models up.

What GRIPS Tests

Native German LLM benchmarks are rare. Most evaluation is English-centric, and German benchmark variants are often created through translation, which quietly smooths away much of what makes German distinctive: its wordplay, idioms, morphology, and letter-level structure. A pun that hinges on German sounds simply does not survive a trip through English.

So we built GRIPS: German Reasoning, Idioms, Puzzles & Wordplay (Sprachspiel), a benchmark built around tasks that require real reasoning and native-level German capabilities to solve. Wordplay, idioms, and letter-level puzzles are the vehicle; reasoning is the point. (Grips is also the German word for wits: „Streng deinen Grips an!“)

GRIPS comes in two parts:

  • Public: an open reference covering examples for all task types, free to inspect and build on. 1,861 questions across 96 mechanics. Includes a hard subset of 119 difficult examples (for models, not necessarily humans).
    Hugging Facehuggingface.co/datasets/ellamind/grips
  • Private: a held-out companion of fresh items, never published, used for the leaderboard so scores reflect genuine generalization rather than exposure to the public benchmark data. To keep it held-out, we only evaluate local models or commercial API models with zero data retention, so prompts are not used for training.

The mechanics span anagrams, hidden words, idioms, Schüttelreime, compound words, and multi-step logic riddles. Every item has a known, checkable answer, programmatically verified, and human reviewed. Try a few yourself: 16 examples straight from the public set, tagged by mechanic:

The GRIPS Leaderboard

We ran a broad field of models on the private held-out set including frontier closed models and open-weight models, reasoning-tuned variants alike. Every answer is graded by an LLM judge against the groundtruth reference answer.

On the GRIPS-hard subset, GPT-5.5 leads at 87.3% in its highest reasoning-effort setting, with Claude Fable 5 close behind at 84.0%: ahead of GPT-5.5’s medium setting (80.7%), but behind its high and extra-high configurations. Below that frontier handful, the picture changes fast: DeepSeek V4 Pro is the strongest open-weight model at 67.3%, already 20 points behind the leader, and the next open-weight model, GLM 5.2, drops to 54.0%, with more DeepSeek and Qwen variants filling the ranks below it. The strongest open-weight model out of the West is Nvidia’s Nemotron 3 Ultra, at 31.3%. Reasoning matters throughout the field, not just at the top: Mistral Medium 3.5 more than doubles its score with reasoning turned on (12.7% → 26.2%), and similar jumps show up across other model families. On the full GRIPS dataset, the field is nearly saturated: the top ten models all clear 90%, and GPT-5.5 tops out at 96.4%, close enough to Claude Fable 5’s 95.5% that their 3.3-point gap on the hard subset shrinks to less than a point here. The closed-vs-open-weight gap narrows just as sharply: GPT-5.5’s 20-point lead over DeepSeek V4 Pro on the hard subset drops to 3.5 points on the full dataset. That’s exactly why the GRIPS-hard subset exists: the full dataset is running out of room to tell top models apart. Reasoning over native German wordplay and letter-level structure still leaves clear headroom at the top.

GRIPS-hard leaderboard: model accuracy on the GRIPS-hard subset

Why Isn’t Fable #1?

Claude Fable 5 currently tops most AI benchmarks. On GRIPS, it only ranks behind GPT-5.5: third on the GRIPS-hard subset, fourth on the full GRIPS dataset. So where did it fail, and for what reason: bad benchmark data, or genuine model mistakes?

We checked every one of its misses, and it turns out to be a mix of genuine model errors and safety refusals. Of course we also ran Fable on the 119-item public GRIPS-hard subset. Here are the items it didn’t get correct:

Five items came back as refusals, all flagged under Anthropic’s “cyber” policy category, and four of those five share one mechanic: short cipher chains encoding unit conversions. That’s more than half of that one mechanic blocked outright, not because the model can’t solve it, but because the input format resembles an exploit string closely enough to trip a safety classifier. The rest of its misses are mostly genuine, sometimes the judge might be a bit too harsh. The failures cluster too: on puns built around a hidden name, Fable often finds a real pun, just not the one being tested, and on puzzles where the underlying rule has to be inferred from only a few examples, it comes up with a elaborate, self-consistent stories or rules, but does not land on the actual solution.

Just Think Harder

The cleanest, most controlled signal comes from holding the model fixed and varying only how much it is allowed to think. On the GRIPS-hard subset, GPT-5.5’s reasoning-effort ladder is monotonic:

Reasoning effortHard subset
xhigh87%
high85%
medium81%
low71%

For most models, test-time scaling buys a lot of performance, and reasoning-tuned models cluster at the top of the board. These puzzles reward step-by-step manipulation (counting letters, testing rearrangements, resolving a pun), not a fast first guess.

Characters, Not Concepts

If we break the scores down by mechanic, a clean pattern appears: models reason well over concepts but stumble over characters. The mechanics they fail most are almost all letter-level manipulation, averaged across every model we ran:

Hardest mechanicAvg. passExample (verbatim)Answer
Spellable from letters56%Welches dieser Wörter lässt sich allein aus den Buchstaben des Namens „KONSTANTIN“ bilden (jeder Buchstabe nur so oft, wie er im Namen vorkommt)? A: TAKT B: OTTO C: RUFA, TAKT
Palindrome58%Welches dieser Wörter liest sich vorwärts wie rückwärts gleich? A: Otmar B: Oskar C: OttoC, Otto
Series from initials63%Diese Buchstaben sind jeweils der erste Buchstabe einer bekannten Reihe. Welcher Buchstabe gehört an die Stelle des Fragezeichens? ? Z D V F S S A N ZE
Hidden name64%Welcher dieser Vornamen versteckt sich – als Buchstabenfolge – im folgenden Satz? A: Felix B: Lukas C: Nora D: Sophie „Vom Gipfel aus genossen wir ein grandioses Panorama.“C, Nora
Middle letters67%Nimmt man von jedem dieser kurzen Wörter nur den mittleren Buchstaben, ergeben sie zusammen ein neues Wort. Welches? ARM, ZEH, OHRREH
Anagram71%Aus genau den Buchstaben von „LAMPE“ lässt sich – in anderer Reihenfolge – ein neues, sinnvolles deutsches Wort bilden: ein Verkehrssignal. Wie lautet es?AMPEL

At the other end, the reasoning mechanics are nearly solved: cyclic position / modular arithmetic (99%), probability (99%), transitive ordering (97%), and arithmetic word problems (90%).

This could just be the tokenization blind spot in plain sight. An LLM sees sub-word tokens, not individual letters, so anything that hinges on spelling, reversing, counting, or rearranging characters is structurally hard, even for frontier models that breeze through multi-step arithmetic. It may also be part of why the Manuel-Neuer pun trips some models: a homophone pun leans on sound and spelling as much as on meaning.

No single item defeats every model, but none of them ace the set either. Even GPT-5.5 at extra-high reasoning effort misses a handful outright, and they are precisely these character-level puzzles, spelling a word from a name’s letters or counting numbers hidden inside words. And yes, we did look at the data to make sure all of GPT-5.5’s failures are genuine.

Try It Yourself

German reasoning and wordplay remain an unsaturated, genuinely hard challenge. Although, frontier commercial models are close. GRIPS gives you a German-native way to measure it, with a fully reproducible evaluation harness.

Run your own model through collect → judge → score on the public benchmark dataset to see where it succeeds or fails.

Acknowledgements

  • This work is supported by the OpenEuroLLM project, co-funded by the Digital Europe Programme under GA no. 101195233.
  • This work is supported by the LLMs4EU project, co-funded by the Digital Europe Programme under GA no. 101198470.
  • This work is supported by the German Federal Ministry for Economic Affairs and Energy (BMWE) through EU-SAI/SOOFI: Sovereign Open Source Foundation Models for European Intelligence (grant number 13IPC040J).
Co-funded by the European UnionFunded by the German Federal Ministry for Economic Affairs and Energy (BMWE)

More articles

Unlock the power of AI

See how our products can help you evaluate, deploy, and monitor AI agents with confidence.