That’s a well-known German pun, and the answer is Neuer, the goalkeeper Manuel Neuer, whose name sounds exactly like the comparative of neu (“newer”). Most Germans get it in a second. Six of the models above land on it too. The rest miss in telling ways: two settle for „manuell“ (right phonetic neighborhood, wrong target), and the others drift to „dabei“ or „älter“: fluent, confident, and wrong. Models rarely say “I don’t know”; they hand back a clean wrong answer instead. And this isn’t a one-off gotcha. Puns are just one category of German puzzles that trip models up.
What GRIPS Tests
Native German LLM benchmarks are rare. Most evaluation is English-centric, and German benchmark variants are often created through translation, which quietly smooths away much of what makes German distinctive: its wordplay, idioms, morphology, and letter-level structure. A pun that hinges on German sounds simply does not survive a trip through English.
So we built GRIPS: German Reasoning, Idioms, Puzzles & Wordplay (Sprachspiel), a benchmark built around tasks that require real reasoning and native-level German capabilities to solve. Wordplay, idioms, and letter-level puzzles are the vehicle; reasoning is the point. (Grips is also the German word for wits: „Streng deinen Grips an!“)
GRIPS comes in two parts:
- Public: an open reference covering examples for all task types, free to inspect and build on. 1,861 questions across 96 mechanics. Includes a hard subset of 119 difficult examples (for models, not necessarily humans).
huggingface.co/datasets/ellamind/grips
- Private: a held-out companion of fresh items, never published, used for the leaderboard so scores reflect genuine generalization rather than exposure to the public benchmark data. To keep it held-out, we only evaluate local models or commercial API models with zero data retention, so prompts are not used for training.
The mechanics span anagrams, hidden words, idioms, Schüttelreime, compound words, and multi-step logic riddles. Every item has a known, checkable answer, programmatically verified, and human reviewed. Try a few yourself: 16 examples straight from the public set, tagged by mechanic:
Aus genau den Buchstaben von „LAMPE“ lässt sich – in anderer Reihenfolge – ein neues, sinnvolles deutsches Wort bilden: ein Verkehrssignal. Wie lautet es?
Welches Wort ergibt vor jedem dieser Wortteile ein sinnvolles zusammengesetztes Wort? -mann -ball -flocke -sturm
Welches dieser Wörter liest sich vorwärts wie rückwärts gleich? A: Ufo B: Uhr C: Uhu
Welcher dieser Vornamen versteckt sich – als Buchstabenfolge – im folgenden Satz? A: Anke B: Felix C: Lukas D: Sophie „Ihm kam ein guter Gedanke in den Sinn.“
Nimmt man von jedem dieser kurzen Wörter nur den mittleren Buchstaben, ergeben sie zusammen ein neues Wort. Welches? UTA, HOF, ARM
Diese Buchstaben sind jeweils der erste Buchstabe einer bekannten Reihe. Welcher Buchstabe gehört an die Stelle des Fragezeichens? F ? H W
Welches dieser Wörter lässt sich allein aus den Buchstaben des Namens „MARTINA“ bilden (jeder Buchstabe nur so oft, wie er im Namen vorkommt)? A: ARM B: QUARZ C: SAU
Bei diesen Wörtern sind nur die inneren Buchstaben verdreht. Lies die Frage und beantworte sie: Wehcler Tag kmmot dkerit ncah Fatierg ?
In dieser Buchstabenfolge stecken zwei ineinander verwobene Wörter; die Reihenfolge der Buchstaben jedes Wortes bleibt dabei erhalten. BAERGPFEL Ein Wort ist BERG. Wie lautet das andere?
Wie viele der sechs Seitenflächen eines Würfels kann man höchstens auf einen Blick sehen?
Man wirft zwei Würfel. Wie wahrscheinlich ist eine Augensumme größer als 9? A: 1/6 B: 1/9 C: 1/4
Vergleiche die Körpergrößen mehrerer Personen: Felix ist 7 cm größer als Anna. Felix ist 8 cm größer als Karla. Mia ist 2 cm größer als Anna. Wer ist am größten?
Schreibe alle Buchstaben des Wortes „MOND“ in alphabetischer Reihenfolge – als eine zusammenhängende Buchstabenfolge.
Bei diesem Code wird „WALD“ zu „VZKC“. Wie lautet nach demselben Muster der Code für „VOGEL“ (fliegt)?
Vervollständige diesen Schüttelreim: „Es klapperten die Klapperschlangen, bis ihre Klappern ___ ___.“
Wie lautet die Zahlenkombination für dieses dreistellige Zahlenschloss? 267: Eine Ziffer ist korrekt, steht aber an der falschen Position. 468: Eine Ziffer ist korrekt, steht aber an der falschen Position. 791: Alle Ziffern sind falsch. 256: Eine Ziffer ist korrekt, steht aber an der falschen Position. 975: Alle Ziffern sind falsch. 590: Eine Ziffer ist korrekt und steht an der richtigen Position. 042: Zwei Ziffern sind korrekt, aber keine steht an der richtigen Position.
The GRIPS Leaderboard
We ran a broad field of models on the private held-out set including frontier closed models and open-weight models, reasoning-tuned variants alike. Every answer is graded by an LLM judge against the groundtruth reference answer.
On the GRIPS-hard subset, GPT-5.5 leads at 87.3% in its highest reasoning-effort setting, with Claude Fable 5 close behind at 84.0%: ahead of GPT-5.5’s medium setting (80.7%), but behind its high and extra-high configurations. Below that frontier handful, the picture changes fast: DeepSeek V4 Pro is the strongest open-weight model at 67.3%, already 20 points behind the leader, and the next open-weight model, GLM 5.2, drops to 54.0%, with more DeepSeek and Qwen variants filling the ranks below it. The strongest open-weight model out of the West is Nvidia’s Nemotron 3 Ultra, at 31.3%. Reasoning matters throughout the field, not just at the top: Mistral Medium 3.5 more than doubles its score with reasoning turned on (12.7% → 26.2%), and similar jumps show up across other model families. On the full GRIPS dataset, the field is nearly saturated: the top ten models all clear 90%, and GPT-5.5 tops out at 96.4%, close enough to Claude Fable 5’s 95.5% that their 3.3-point gap on the hard subset shrinks to less than a point here. The closed-vs-open-weight gap narrows just as sharply: GPT-5.5’s 20-point lead over DeepSeek V4 Pro on the hard subset drops to 3.5 points on the full dataset. That’s exactly why the GRIPS-hard subset exists: the full dataset is running out of room to tell top models apart. Reasoning over native German wordplay and letter-level structure still leaves clear headroom at the top.
Why Isn’t Fable #1?
Claude Fable 5 currently tops most AI benchmarks. On GRIPS, it only ranks behind GPT-5.5: third on the GRIPS-hard subset, fourth on the full GRIPS dataset. So where did it fail, and for what reason: bad benchmark data, or genuine model mistakes?
We checked every one of its misses, and it turns out to be a mix of genuine model errors and safety refusals. Of course we also ran Fable on the 119-item public GRIPS-hard subset. Here are the items it didn’t get correct:
Vervollständige dieses bekannte Wortspiel: „Egal wie viele Kühe du ihm hinstellst, Bastian ist ___.“
Vervollständige dieses bekannte Wortspiel: „Egal wie kalt es draußen ist, Leonardo fährt ___.“
Vervollständige dieses bekannte Wortspiel: „Egal wie laut du Bach hörst, Heiner hört ___.“
Nach einer bestimmten Logik gilt: Fluss = 19; Tal = 13; Moor = 20. Welchen Wert hat dann „See“?
Vervollständige dieses bekannte Wortspiel: „Egal wie scharf du bist, der Sänger ist ___.“
Vervollständige diesen Flachwitz: „Ich habe einen Schornsteinfeger angerufen, aber …“ A: er war auf dem Dach B: er war besetzt C: er war verrußt D: er hatte abgekehrt
Scherzfrage: Was ist schwarz-weiß und kann nicht um die Ecke schauen? A: ein Zebra B: ein Pinguin C: eine Zeitung D: ein Panda
Vervollständige die absichtlich verdrehte Redewendung (ein Kalauer): „Wer anderen eine Grube gräbt, …“ A: hat ein Grubengerät. B: ist selbst hineingefallen. C: sollte Bauarbeiter werden. D: kennt sich mit Erde aus.
Scherzfrage: Warum können Fische so gut rechnen? A: Weil sie im Schwarm denken B: Weil sie eine Kiemen-Logik haben C: Weil sie sich gut mit Wurzeln auskennen D: Weil sie im Meer viele Algen haben
Wie nennt man – mit einem englisch-deutschen Wortspiel – eine Katze, die im Internet surft? A: eine Web-Katze B: eine Catze C: eine Net-Mieze D: eine Cat-Maus
Vervollständige dieses bekannte Wortspiel: „Egal wie cool du bist, Karl ist ___.“
Vervollständige dieses Wortspiel (das gesuchte Wort ist ein Name oder klingt wie einer): „Egal wie gut du fährst, Züge fahren ___.“
Vervollständige dieses bekannte Wortspiel: „Egal wie schwer du es hast, Atlas trägt ___.“
Bei diesen Zahlen sind die Silben durcheinandergeraten, doch jede lässt sich zu genau einer sinnvollen Zahl zusammensetzen. Welche ist die kleinste? A: FÜNFUNDZIGSIEB B: ACHTUNDZIGZWEI C: ZIGSECHSNEUNUND
Vervollständige diesen Schüttelreim: „Es gibt so viele stumme Denker, doch häufiger sind ___ ___.“
Nach einer bestimmten Logik gilt: See = 13; Berg = 14; Wald = 14. Welchen Wert hat dann „Wiese“?
In dieser verschlüsselten Kette steht jeweils „‹Anzahl›‹Anfangsbuchstabe›s 1 ‹nächster Buchstabe›“. Welcher Buchstabe gehört an die Stelle des Fragezeichens? 12?s1D12Ds1G
Five items came back as refusals, all flagged under Anthropic’s “cyber” policy category, and four of those five share one mechanic: short cipher chains encoding unit conversions. That’s more than half of that one mechanic blocked outright, not because the model can’t solve it, but because the input format resembles an exploit string closely enough to trip a safety classifier. The rest of its misses are mostly genuine, sometimes the judge might be a bit too harsh. The failures cluster too: on puns built around a hidden name, Fable often finds a real pun, just not the one being tested, and on puzzles where the underlying rule has to be inferred from only a few examples, it comes up with a elaborate, self-consistent stories or rules, but does not land on the actual solution.
Just Think Harder
The cleanest, most controlled signal comes from holding the model fixed and varying only how much it is allowed to think. On the GRIPS-hard subset, GPT-5.5’s reasoning-effort ladder is monotonic:
| Reasoning effort | Hard subset |
|---|---|
| xhigh | 87% |
| high | 85% |
| medium | 81% |
| low | 71% |
For most models, test-time scaling buys a lot of performance, and reasoning-tuned models cluster at the top of the board. These puzzles reward step-by-step manipulation (counting letters, testing rearrangements, resolving a pun), not a fast first guess.
Characters, Not Concepts
If we break the scores down by mechanic, a clean pattern appears: models reason well over concepts but stumble over characters. The mechanics they fail most are almost all letter-level manipulation, averaged across every model we ran:
| Hardest mechanic | Avg. pass | Example (verbatim) | Answer |
|---|---|---|---|
| Spellable from letters | 56% | Welches dieser Wörter lässt sich allein aus den Buchstaben des Namens „KONSTANTIN“ bilden (jeder Buchstabe nur so oft, wie er im Namen vorkommt)? A: TAKT B: OTTO C: RUF | A, TAKT |
| Palindrome | 58% | Welches dieser Wörter liest sich vorwärts wie rückwärts gleich? A: Otmar B: Oskar C: Otto | C, Otto |
| Series from initials | 63% | Diese Buchstaben sind jeweils der erste Buchstabe einer bekannten Reihe. Welcher Buchstabe gehört an die Stelle des Fragezeichens? ? Z D V F S S A N Z | E |
| Hidden name | 64% | Welcher dieser Vornamen versteckt sich – als Buchstabenfolge – im folgenden Satz? A: Felix B: Lukas C: Nora D: Sophie „Vom Gipfel aus genossen wir ein grandioses Panorama.“ | C, Nora |
| Middle letters | 67% | Nimmt man von jedem dieser kurzen Wörter nur den mittleren Buchstaben, ergeben sie zusammen ein neues Wort. Welches? ARM, ZEH, OHR | REH |
| Anagram | 71% | Aus genau den Buchstaben von „LAMPE“ lässt sich – in anderer Reihenfolge – ein neues, sinnvolles deutsches Wort bilden: ein Verkehrssignal. Wie lautet es? | AMPEL |
At the other end, the reasoning mechanics are nearly solved: cyclic position / modular arithmetic (99%), probability (99%), transitive ordering (97%), and arithmetic word problems (90%).
This could just be the tokenization blind spot in plain sight. An LLM sees sub-word tokens, not individual letters, so anything that hinges on spelling, reversing, counting, or rearranging characters is structurally hard, even for frontier models that breeze through multi-step arithmetic. It may also be part of why the Manuel-Neuer pun trips some models: a homophone pun leans on sound and spelling as much as on meaning.
No single item defeats every model, but none of them ace the set either. Even GPT-5.5 at extra-high reasoning effort misses a handful outright, and they are precisely these character-level puzzles, spelling a word from a name’s letters or counting numbers hidden inside words. And yes, we did look at the data to make sure all of GPT-5.5’s failures are genuine.
Try It Yourself
German reasoning and wordplay remain an unsaturated, genuinely hard challenge. Although, frontier commercial models are close. GRIPS gives you a German-native way to measure it, with a fully reproducible evaluation harness.
- Dataset:
huggingface.co/datasets/ellamind/grips
- Evaluation code: github.com/ellamind/grips-eval
Run your own model through collect → judge → score on the public benchmark dataset to see where it succeeds or fails.
Acknowledgements
- This work is supported by the OpenEuroLLM project, co-funded by the Digital Europe Programme under GA no. 101195233.
- This work is supported by the LLMs4EU project, co-funded by the Digital Europe Programme under GA no. 101198470.
- This work is supported by the German Federal Ministry for Economic Affairs and Energy (BMWE) through EU-SAI/SOOFI: Sovereign Open Source Foundation Models for European Intelligence (grant number 13IPC040J).