Test · AI feedback · Writing B2
Published · October 6, 2026 · 9 min read

ChatGPT vs GoetheCoach for the Goethe Exam: We Graded the Same B2 Text 5 Times

ChatGPT is a strong study partner for the Goethe exam: it explains grammar, simulates Sprechen by voice and drafts practice tasks. As an examiner it is less reliable. In our test it gave the same B2 text 82 to 89 points and said all Leitpunkte were covered while one was missing. GoetheCoach flagged the gap in every run.

Split image: on the left a chat bubble with shifting scores 82 and 89, on the right a GoetheCoach checklist of four Leitpunkte with one marked red as missing

Disclosure: GoetheCoach is our product, so read this as a vendor test. We publish the test text, the prompt and every result below so you can repeat it yourself. GoetheCoach and this article are not affiliated with the Goethe-Institut or OpenAI.

The short answer

You want to…Better choiceWhy
Understand a grammar rule or a wordChatGPTPatient, fast, explains in any language
Practise speaking (Sprechen) out loudChatGPT (voice mode)GoetheCoach covers Schreiben only
Get practice tasks and model textsBothChatGPT drafts freely; GoetheCoach uses the official task formats
Know whether your Schreiben text would passGoetheCoachChecks every Leitpunkt, counts words in code, scores by a fixed rubric
Get the same score for the same text twiceGoetheCoachScore is computed by code; identical resubmissions return the identical result
Get a real exam gradeNeitherOnly Goethe-Institut examiners grade the exam. Both give practice estimates

How we tested

On 6 October 2026 we wrote one B2 Forumsbeitrag (Schreiben Teil 1) with three properties a real examiner would notice:

The full test text:

In den letzten Jahren arbeiten immer mehr Menschen von zu Hause aus. Meiner Meinung nach sollte Homeoffice für alle Berufsgruppen dauerhaft möglich sein, soweit es die Arbeit erlaubt. Zuerst spart man viel Zeit, weil man nicht jeden Tag pendeln muss. Ich selbst habe früher fast zwei Stunden pro Tag im Zug verbracht. Diese Zeit kann man jetzt für die Familie oder für Sport nutzen. Außerdem ist man zu Hause oft konzentrierter, da es keine Kollegen gibt, die einem ständig stören. Ein weiteres Argument ist die Umwelt: Wenn weniger Leute mit dem Auto zur Arbeit fahren, gibt es weniger Stau und weniger Abgase in den Städten. Natürlich sagen manche, dass man im Homeoffice einsam wird und der Kontakt zum Team verloren geht. Das stimmt teilweise. Aber man kann sich regelmäßig in Videokonferenzen austauschen, und gute Teams finden auch online Wege, um in Kontakt zu bleiben. Zusammenfassend bin ich überzeugt, dass die Vorteile deutlich überwiegen. Firmen sollten ihren Mitarbeiter mehr Vertrauen schenken und ihnen die Freiheit geben, selbst zu entscheiden, wo sie am besten arbeiten.

The task: "Homeoffice sollte für alle Berufsgruppen dauerhaft möglich sein." with four Leitpunkte: argue for or against, explain arguments, deal with a counter-argument, name alternatives. This is the most common way B2 texts lose points: in our Leitpunkte study, Leitpunkt 4 was missing in 46% of B2 texts.

ChatGPT: chatgpt.com, Free plan, default model, a fresh Temporary chat each time (no memory, no custom instructions), the same German prompt five times: "Please grade my text like a Goethe examiner: give me a score from 0 to 100 and tell me whether I would have passed (pass mark 60)." followed by the task, the four Leitpunkte and the text.

GoetheCoach: the free demo grader on goethecoach.de (no account), five runs with the same task and text. The demo is our lightest version: a smaller model and no result cache.

Results: 5 runs each

RunChatGPT totalChatGPT task scoreSays all 4 Leitpunkte covered?GoetheCoach totalGoetheCoach task scoreLeitpunkt 4 flagged as missing?
18223/25Yes8219/25Yes
28222/25Yes7419/25Yes
38922/25Yes7419/25Yes
48222/25Yes7419/25Yes
58222/25Yes7419/25Yes

Both tools said the text passes, and both varied: ChatGPT by 7 points, the GoetheCoach demo by 8. So "AI is inconsistent, we are not" would be the wrong headline. The difference is where the variation comes from and what the feedback tells you.

ChatGPT vs GoetheCoach: the same B2 text, graded 5 times each (1:28). Test from October 6, 2026. Not affiliated with the Goethe-Institut or OpenAI.
Video transcript

Can ChatGPT grade your Goethe exam? We tested it against GoetheCoach, with the same text, five times each. Our test text is a B2 forum post: one hundred seventy-three words, with two grammar mistakes. And one trap: Leitpunkt four, name alternatives, is completely missing. We pasted it into five fresh ChatGPT chats. Eighty-two. Eighty-two. Eighty-nine. Eighty-two. Eighty-two. A pass, every time. And every time it said: all four Leitpunkte are covered. A few lines later, the same answer admits that the alternatives are missing. And the word count? Its guesses were 190, 155 and 170 words. GoetheCoach flagged Leitpunkt four as missing in all five runs. The same task score every time, and the words are counted, not guessed. The difference: the AI only reports what it sees, with a quote for every Leitpunkt. The points are calculated by code, with a fixed rubric, like in the real exam. To be fair: for grammar and speaking practice, ChatGPT is a great tutor. But to know if your text would pass, you need an examiner, not a cheerleader. Read the full test and check your own text on goethecoach dot de.

What went wrong in ChatGPT's feedback

1. It contradicted itself on the Leitpunkte, every time

In all five runs ChatGPT wrote that all four Leitpunkte were covered ("grundsätzlich behandelt"), and later in the same answer that the text names no real alternatives. Run 3: "Du behandelst alle vier Leitpunkte", then "Alternativen sind allerdings nur sehr indirekt vorhanden." The task score stayed at 22–23 of 25, as if a quarter of the task were a detail.

The official B2 model test says examiners look at "wie genau die Inhaltspunkte bearbeitet sind": how precisely the content points are handled. For a learner this is the costly part. You read "all Leitpunkte covered, 82 points, clear pass" at the top and stop worrying about the one thing that costs most points in the real exam.

2. It guessed the word count

ChatGPT estimated the length three times: about 190, about 155 and about 170 words. The text has 173. At B2 the minimum is 150 words; an estimate of 155 tells you that you are on the edge, an estimate of 190 tells you that you are safe. A language model reads text in tokens, not words, so counting is a guess unless it runs code.

3. It caught different errors in different runs

die einem stören was found in all five runs. ihren Mitarbeiter was called a grammar error in three runs, a style issue in one ("verständlich, aber stilistisch besser") and was not mentioned at all in one. Same text, same prompt.

4. It hedged its own number

Run 4 gave 82/100 and then added that a stricter examiner would give 75–80 and a lenient one over 80. That is honest, but it shows the number is a feeling, not a calculation.

5. Exam facts: mostly right, one invented detail

We also asked how B2 Schreiben is structured. With web search on, ChatGPT got the essentials right: two parts, 75 minutes, a forum post of at least 150 words, a formal message of at least 100 words, 60 of 100 points to pass. It also stated a fixed split of 60 points for Teil 1 and 40 for Teil 2 and attributed it to the Goethe-Institut. The official B2 model test does not publish a per-part split, and third-party sources disagree. Small detail, but it is exactly the kind of confident claim you cannot check from inside the chat.

What ChatGPT did well

The language feedback itself was good. Every run explained einen vs einem correctly, suggested natural B2 phrases (Ein weiterer Vorteil besteht darin, dass …) and gave a concrete example of an alternative (hybrides Arbeitsmodell). If you treat it as a tutor rather than an examiner, it is useful.

What GoetheCoach does differently

GoetheCoach also uses a language model. The difference is the architecture around it:

  1. The model does not give points. It reports observations: for each Leitpunkt full / partial / absent with the sentence that covers it, each error with a quote, and a band for Kohärenz, Wortschatz and Strukturen.
  2. Code turns observations into the score. The same observations always give the same score. That is why the task score was 19/25 in all five runs: "3 complete, 0 partial, 1 missing".
  3. Words are counted in code, never estimated.
  4. Temperature 0 and a result cache in the app: if you submit the identical text again, you get the identical result, not a new roll of the dice.
  5. The rubric is versioned (rubric 1.0) and built on the four official criteria: Aufgabenerfüllung, Kohärenz, Wortschatz, Strukturen. See how Goethe writing is scored.

This is also how the real exam works. According to the Goethe-Institut's implementation rules for B2, Schreiben is graded by two independent raters against fixed criteria, and only the predefined point values per criterion may be awarded, with no values in between. A fixed procedure is the point.

It also matches what researchers found makes AI grading stable. A 2025 study in Education Sciences (Garcia-Varela et al.) found that ChatGPT gave volatile grades with minimal guidance; a rubric reduced the variation but did not remove it. Grading only became consistent with an example-rich rubric, a fixed output format, low temperature and a decision table that turns judgements into points.

Where GoetheCoach is weaker

What research says about ChatGPT as a grader

StudyWhat was testedFinding
Seßler et al., LAK 202520 German school essays, 37 teachers, GPT-3.5, GPT-4, o1 and open modelso1 correlated well with teachers (ρ = .74), but models tended to give higher scores and were weaker on content quality
Kubesch, Huber & Havas, 2026101 Austrian Matura German essays, four open-weight models with the official rubricAt most 40.6% agreement on rubric sub-dimensions; only 32.8% of final grades matched the expert
Lau, 2026GPT-4o, Gemini and Claude models as judges, repeated runsSame input, different scores, even at temperature 0
Saricaoglu & Bilki, 2025ChatGPT-4 feedback on 35 B2 English essaysError identification was 96–99% accurate, but only 37% of task-response feedback was specific
Yancey et al., BEA 2023GPT-4 rating short L2 English essays on the CEFR scaleWith calibration examples, close to dedicated scoring systems; agreement varied by the writer's first language

Two patterns repeat. Language models are good at spotting language errors and weaker at judging whether the task was done. And their scores drift unless a fixed procedure turns judgements into points. Both are exactly what costs points in the Goethe Schreiben module, where task fulfilment is one of four criteria.

There is also a tone problem. In April 2025 OpenAI rolled back a ChatGPT update because the model had become sycophantic: too agreeable and too quick to praise. That has been fixed, but a general assistant is tuned to be helpful and encouraging. An examiner is tuned to find what is missing.

How to use ChatGPT well for Goethe prep

If you use ChatGPT, these habits close most of the gaps we saw:

  1. Ask for a Leitpunkt table first, score second. "For each Leitpunkt, quote the sentence that covers it, or write MISSING." A forced quote makes it harder to say "covered" when nothing is there.
  2. Count words yourself. Any editor does it in one click.
  3. Ask for errors as a list, not a grade. Error spotting is ChatGPT's strength; scoring is its weakness.
  4. Run it twice. If the two scores differ, neither is reliable.
  5. Use it for what it is great at: explanations, vocabulary, Redemittel, and spoken practice for Sprechen.
  6. Check exam facts on goethe.de. Formats, times and pass marks are published there.

For a quick check of a finished text against the rubric, try the free Goethe Writing Check or practise a full task with GoetheCoach.

FAQ

Can ChatGPT grade my Goethe Schreiben text reliably?

Not reliably. In our test it gave the same B2 text 82 to 89 points across five chats and claimed all four Leitpunkte were covered although one was missing. Use it for explanations and error spotting, and check task coverage yourself.

Is ChatGPT good for Goethe exam preparation?

Yes, as a tutor. It explains grammar, suggests phrases, generates practice texts and its voice mode is useful for Sprechen. Be careful with scores, word counts and exam facts it states without a source.

Is GoetheCoach better than ChatGPT?

For grading Schreiben, yes: it checks each Leitpunkt, counts words in code and computes the score from a fixed rubric. For speaking practice and open questions, ChatGPT is more flexible. Many learners use both.

Does GoetheCoach use ChatGPT?

GoetheCoach uses a language model to read the text, but the model does not award points. It reports observations, and code computes the score from them using a versioned rubric based on the four official Goethe criteria.

Will the Goethe exam give me the same score as an AI tool?

Not necessarily. Only Goethe-Institut examiners grade the exam: Schreiben is rated by two independent raters, and a third rater decides if they disagree around the pass mark. Every AI score, including GoetheCoach's, is a practice estimate.

Which ChatGPT plan did you test?

The free plan on chatgpt.com with the default model, in Temporary chats without memory, on 6 October 2026. Paid plans and reasoning modes may behave differently, so repeat the test with your own text.

Sources

Get every Leitpunkt checked, not guessed

Write a B2 or C1 task and see which Leitpunkte you covered, your exact word count and why you got each point.

Start free