Disclosure: GoetheCoach is our product, so read this as a vendor test. We publish the test text, the prompt and every result below so you can repeat it yourself. GoetheCoach and this article are not affiliated with the Goethe-Institut or OpenAI.
The short answer
| You want to… | Better choice | Why |
|---|---|---|
| Understand a grammar rule or a word | ChatGPT | Patient, fast, explains in any language |
| Practise speaking (Sprechen) out loud | ChatGPT (voice mode) | GoetheCoach covers Schreiben only |
| Get practice tasks and model texts | Both | ChatGPT drafts freely; GoetheCoach uses the official task formats |
| Know whether your Schreiben text would pass | GoetheCoach | Checks every Leitpunkt, counts words in code, scores by a fixed rubric |
| Get the same score for the same text twice | GoetheCoach | Score is computed by code; identical resubmissions return the identical result |
| Get a real exam grade | Neither | Only Goethe-Institut examiners grade the exam. Both give practice estimates |
How we tested
On 6 October 2026 we wrote one B2 Forumsbeitrag (Schreiben Teil 1) with three properties a real examiner would notice:
- 173 words, so above the minimum of 150.
- Leitpunkt 4 is missing. The task asks for Alternativen; the text names none. It only answers the counter-argument (loneliness) with video calls.
- Two grammar errors: die einem ständig stören (should be einen) and ihren Mitarbeiter mehr Vertrauen schenken (dative plural: Mitarbeitern).
The full test text:
In den letzten Jahren arbeiten immer mehr Menschen von zu Hause aus. Meiner Meinung nach sollte Homeoffice für alle Berufsgruppen dauerhaft möglich sein, soweit es die Arbeit erlaubt. Zuerst spart man viel Zeit, weil man nicht jeden Tag pendeln muss. Ich selbst habe früher fast zwei Stunden pro Tag im Zug verbracht. Diese Zeit kann man jetzt für die Familie oder für Sport nutzen. Außerdem ist man zu Hause oft konzentrierter, da es keine Kollegen gibt, die einem ständig stören. Ein weiteres Argument ist die Umwelt: Wenn weniger Leute mit dem Auto zur Arbeit fahren, gibt es weniger Stau und weniger Abgase in den Städten. Natürlich sagen manche, dass man im Homeoffice einsam wird und der Kontakt zum Team verloren geht. Das stimmt teilweise. Aber man kann sich regelmäßig in Videokonferenzen austauschen, und gute Teams finden auch online Wege, um in Kontakt zu bleiben. Zusammenfassend bin ich überzeugt, dass die Vorteile deutlich überwiegen. Firmen sollten ihren Mitarbeiter mehr Vertrauen schenken und ihnen die Freiheit geben, selbst zu entscheiden, wo sie am besten arbeiten.
The task: "Homeoffice sollte für alle Berufsgruppen dauerhaft möglich sein." with four Leitpunkte: argue for or against, explain arguments, deal with a counter-argument, name alternatives. This is the most common way B2 texts lose points: in our Leitpunkte study, Leitpunkt 4 was missing in 46% of B2 texts.
ChatGPT: chatgpt.com, Free plan, default model, a fresh Temporary chat each time (no memory, no custom instructions), the same German prompt five times: "Please grade my text like a Goethe examiner: give me a score from 0 to 100 and tell me whether I would have passed (pass mark 60)." followed by the task, the four Leitpunkte and the text.
GoetheCoach: the free demo grader on goethecoach.de (no account), five runs with the same task and text. The demo is our lightest version: a smaller model and no result cache.
Results: 5 runs each
| Run | ChatGPT total | ChatGPT task score | Says all 4 Leitpunkte covered? | GoetheCoach total | GoetheCoach task score | Leitpunkt 4 flagged as missing? |
|---|---|---|---|---|---|---|
| 1 | 82 | 23/25 | Yes | 82 | 19/25 | Yes |
| 2 | 82 | 22/25 | Yes | 74 | 19/25 | Yes |
| 3 | 89 | 22/25 | Yes | 74 | 19/25 | Yes |
| 4 | 82 | 22/25 | Yes | 74 | 19/25 | Yes |
| 5 | 82 | 22/25 | Yes | 74 | 19/25 | Yes |
Both tools said the text passes, and both varied: ChatGPT by 7 points, the GoetheCoach demo by 8. So "AI is inconsistent, we are not" would be the wrong headline. The difference is where the variation comes from and what the feedback tells you.
Video transcript
Can ChatGPT grade your Goethe exam? We tested it against GoetheCoach, with the same text, five times each. Our test text is a B2 forum post: one hundred seventy-three words, with two grammar mistakes. And one trap: Leitpunkt four, name alternatives, is completely missing. We pasted it into five fresh ChatGPT chats. Eighty-two. Eighty-two. Eighty-nine. Eighty-two. Eighty-two. A pass, every time. And every time it said: all four Leitpunkte are covered. A few lines later, the same answer admits that the alternatives are missing. And the word count? Its guesses were 190, 155 and 170 words. GoetheCoach flagged Leitpunkt four as missing in all five runs. The same task score every time, and the words are counted, not guessed. The difference: the AI only reports what it sees, with a quote for every Leitpunkt. The points are calculated by code, with a fixed rubric, like in the real exam. To be fair: for grammar and speaking practice, ChatGPT is a great tutor. But to know if your text would pass, you need an examiner, not a cheerleader. Read the full test and check your own text on goethecoach dot de.
What went wrong in ChatGPT's feedback
1. It contradicted itself on the Leitpunkte, every time
In all five runs ChatGPT wrote that all four Leitpunkte were covered ("grundsätzlich behandelt"), and later in the same answer that the text names no real alternatives. Run 3: "Du behandelst alle vier Leitpunkte", then "Alternativen sind allerdings nur sehr indirekt vorhanden." The task score stayed at 22–23 of 25, as if a quarter of the task were a detail.
The official B2 model test says examiners look at "wie genau die Inhaltspunkte bearbeitet sind": how precisely the content points are handled. For a learner this is the costly part. You read "all Leitpunkte covered, 82 points, clear pass" at the top and stop worrying about the one thing that costs most points in the real exam.
2. It guessed the word count
ChatGPT estimated the length three times: about 190, about 155 and about 170 words. The text has 173. At B2 the minimum is 150 words; an estimate of 155 tells you that you are on the edge, an estimate of 190 tells you that you are safe. A language model reads text in tokens, not words, so counting is a guess unless it runs code.
3. It caught different errors in different runs
die einem stören was found in all five runs. ihren Mitarbeiter was called a grammar error in three runs, a style issue in one ("verständlich, aber stilistisch besser") and was not mentioned at all in one. Same text, same prompt.
4. It hedged its own number
Run 4 gave 82/100 and then added that a stricter examiner would give 75–80 and a lenient one over 80. That is honest, but it shows the number is a feeling, not a calculation.
5. Exam facts: mostly right, one invented detail
We also asked how B2 Schreiben is structured. With web search on, ChatGPT got the essentials right: two parts, 75 minutes, a forum post of at least 150 words, a formal message of at least 100 words, 60 of 100 points to pass. It also stated a fixed split of 60 points for Teil 1 and 40 for Teil 2 and attributed it to the Goethe-Institut. The official B2 model test does not publish a per-part split, and third-party sources disagree. Small detail, but it is exactly the kind of confident claim you cannot check from inside the chat.
What ChatGPT did well
The language feedback itself was good. Every run explained einen vs einem correctly, suggested natural B2 phrases (Ein weiterer Vorteil besteht darin, dass …) and gave a concrete example of an alternative (hybrides Arbeitsmodell). If you treat it as a tutor rather than an examiner, it is useful.
What GoetheCoach does differently
GoetheCoach also uses a language model. The difference is the architecture around it:
- The model does not give points. It reports observations: for each Leitpunkt full / partial / absent with the sentence that covers it, each error with a quote, and a band for Kohärenz, Wortschatz and Strukturen.
- Code turns observations into the score. The same observations always give the same score. That is why the task score was 19/25 in all five runs: "3 complete, 0 partial, 1 missing".
- Words are counted in code, never estimated.
- Temperature 0 and a result cache in the app: if you submit the identical text again, you get the identical result, not a new roll of the dice.
- The rubric is versioned (rubric 1.0) and built on the four official criteria: Aufgabenerfüllung, Kohärenz, Wortschatz, Strukturen. See how Goethe writing is scored.
This is also how the real exam works. According to the Goethe-Institut's implementation rules for B2, Schreiben is graded by two independent raters against fixed criteria, and only the predefined point values per criterion may be awarded, with no values in between. A fixed procedure is the point.
It also matches what researchers found makes AI grading stable. A 2025 study in Education Sciences (Garcia-Varela et al.) found that ChatGPT gave volatile grades with minimal guidance; a rubric reduced the variation but did not remove it. Grading only became consistent with an example-rich rubric, a fixed output format, low temperature and a decision table that turns judgements into points.
Where GoetheCoach is weaker
- The demo also varied (74 vs 82). The variation came from one judgement on vocabulary and grammar bands, not from the Leitpunkte. The signed-in app uses a larger model and caches results, but a band judgement can still differ between two different texts of similar quality.
- Only Schreiben. For Sprechen, Lesen and Hören you need other material, and ChatGPT's voice mode is genuinely useful for speaking practice.
- Not an official grade. No tool can tell you your exact exam result. Treat every score as a practice estimate; for a second opinion, GoetheCoach offers a teacher check.
What research says about ChatGPT as a grader
| Study | What was tested | Finding |
|---|---|---|
| Seßler et al., LAK 2025 | 20 German school essays, 37 teachers, GPT-3.5, GPT-4, o1 and open models | o1 correlated well with teachers (ρ = .74), but models tended to give higher scores and were weaker on content quality |
| Kubesch, Huber & Havas, 2026 | 101 Austrian Matura German essays, four open-weight models with the official rubric | At most 40.6% agreement on rubric sub-dimensions; only 32.8% of final grades matched the expert |
| Lau, 2026 | GPT-4o, Gemini and Claude models as judges, repeated runs | Same input, different scores, even at temperature 0 |
| Saricaoglu & Bilki, 2025 | ChatGPT-4 feedback on 35 B2 English essays | Error identification was 96–99% accurate, but only 37% of task-response feedback was specific |
| Yancey et al., BEA 2023 | GPT-4 rating short L2 English essays on the CEFR scale | With calibration examples, close to dedicated scoring systems; agreement varied by the writer's first language |
Two patterns repeat. Language models are good at spotting language errors and weaker at judging whether the task was done. And their scores drift unless a fixed procedure turns judgements into points. Both are exactly what costs points in the Goethe Schreiben module, where task fulfilment is one of four criteria.
There is also a tone problem. In April 2025 OpenAI rolled back a ChatGPT update because the model had become sycophantic: too agreeable and too quick to praise. That has been fixed, but a general assistant is tuned to be helpful and encouraging. An examiner is tuned to find what is missing.
How to use ChatGPT well for Goethe prep
If you use ChatGPT, these habits close most of the gaps we saw:
- Ask for a Leitpunkt table first, score second. "For each Leitpunkt, quote the sentence that covers it, or write MISSING." A forced quote makes it harder to say "covered" when nothing is there.
- Count words yourself. Any editor does it in one click.
- Ask for errors as a list, not a grade. Error spotting is ChatGPT's strength; scoring is its weakness.
- Run it twice. If the two scores differ, neither is reliable.
- Use it for what it is great at: explanations, vocabulary, Redemittel, and spoken practice for Sprechen.
- Check exam facts on goethe.de. Formats, times and pass marks are published there.
For a quick check of a finished text against the rubric, try the free Goethe Writing Check or practise a full task with GoetheCoach.
FAQ
Can ChatGPT grade my Goethe Schreiben text reliably?
Not reliably. In our test it gave the same B2 text 82 to 89 points across five chats and claimed all four Leitpunkte were covered although one was missing. Use it for explanations and error spotting, and check task coverage yourself.
Is ChatGPT good for Goethe exam preparation?
Yes, as a tutor. It explains grammar, suggests phrases, generates practice texts and its voice mode is useful for Sprechen. Be careful with scores, word counts and exam facts it states without a source.
Is GoetheCoach better than ChatGPT?
For grading Schreiben, yes: it checks each Leitpunkt, counts words in code and computes the score from a fixed rubric. For speaking practice and open questions, ChatGPT is more flexible. Many learners use both.
Does GoetheCoach use ChatGPT?
GoetheCoach uses a language model to read the text, but the model does not award points. It reports observations, and code computes the score from them using a versioned rubric based on the four official Goethe criteria.
Will the Goethe exam give me the same score as an AI tool?
Not necessarily. Only Goethe-Institut examiners grade the exam: Schreiben is rated by two independent raters, and a third rater decides if they disagree around the pass mark. Every AI score, including GoetheCoach's, is a practice estimate.
Which ChatGPT plan did you test?
The free plan on chatgpt.com with the default model, in Temporary chats without memory, on 6 October 2026. Paid plans and reasoning modes may behave differently, so repeat the test with your own text.
Sources
- Goethe-Institut: Goethe-Zertifikat B2 model test, Schreiben
- Goethe-Institut: Durchführungsbestimmungen Goethe-Zertifikat B2 (§ 4.3 Modul Schreiben)
- Seßler, Fürstenberg, Bühler & Kasneci (2025): Can AI grade your essays?, LAK '25
- Kubesch, Huber & Havas (2026): Evaluating Austrian A-Level German Essays with Large Language Models
- Lau (2026): Same Input, Different Scores
- Garcia-Varela et al. (2025): ChatGPT as a Stable and Fair Tool for Automated Essay Scoring, Education Sciences 15(8)
- Saricaoglu & Bilki (2025): The capacity of ChatGPT-4 for L2 writing assessment, Annual Review of Applied Linguistics
- Yancey, Laflair, Verardi & Burstein (2023): Rating Short L2 Essays on the CEFR Scale with GPT-4, BEA 2023
- OpenAI (2025): Expanding on what we missed with sycophancy
Get every Leitpunkt checked, not guessed
Write a B2 or C1 task and see which Leitpunkte you covered, your exact word count and why you got each point.
Start free