What to take away
- A general chatbot is good at maths drills, generating cases and critiquing a written structure, and weak at scoring you truthfully without a strict prompt
- The most common complaint about AI case tools is feedback that is generic and too forgiving, and you can test for it in two minutes with a deliberately bad answer
- Useful AI feedback quotes what you said, names the skill, separates a wrong number from a missing sense-check, and gives a score that moves when you improve
- Marking against a fixed first-round bar of 3.5, not a curve, is what lets a weak case score 2 and a strong one score 4
- Spoken practice catches hedging, pace and answer order that typed practice hides, and neither replaces a human mock before a real round
- Use AI for reps in the early weeks and spend your scarce human sessions on reading a room and a partner's style
What counts as AI case interview practice in 2026?
Three different things get called AI practice, and candidates who try the wrong one first often decide the whole idea does not work. The first is a general chatbot you steer yourself: you open ChatGPT, Claude or Gemini, paste in a case, and ask the model to play the interviewer. The second is a purpose-built text interviewer, where a fixed case and a fixed rubric sit behind the chat window and you get a scored report at the end. The third is a purpose-built voice interviewer that runs a spoken case start to finish, the way a first round is run.
The typed-versus-spoken split deserves more weight than it usually gets. A case interview is marked on what comes out of your mouth under time, with an interviewer cutting across you now and then. In a typed chat you can stop for two minutes to think, rewrite a sentence three times and delete the hedge before anyone sees it. None of that is allowed in the room, so typed practice trains a slightly different skill.
Both ChatGPT and Claude now have voice modes, so a general chatbot no longer has to be typed. That closes part of the gap. It leaves open the questions this guide is really about: whether the tool marks you against a standard, whether it tells you the truth when you are bad, and whether the case it runs stays consistent from minute one to minute thirty.
We build one of these tools, and we say so throughout. Where we describe how our own marking works, treat it as a claim from the people selling it and test it the same way we suggest you test everyone else's. The coaches on MBB Ready are AI personas, and how case interviews are scored sets out the rubric they mark against in full.
Does practising case interviews with ChatGPT help?
Yes, for some parts of preparation, and the useful parts are narrower than the marketing around 'AI mock interviews' suggests. Posters on PrepLounge's forum, several of them coaches, list roughly the same uses in a February 2025 thread on using ChatGPT for case practice: brainstorming with a structure you supplied, industry background, maths problems, tightening phrasing and generating practice questions. The same thread is just as consistent about the failures, which they describe as superficial frameworks, average insights and dependence on what you feed in. Those are candidate and coach reports rather than measurements, and they match what we see in the transcripts we mark.
PrepLounge's own article on AI for interview prep, published in July 2025, draws the same line from the platform's side. It lists frameworking, maths practice, brainstorming, business judgement and fit questions as places AI helps, and it says plainly that ChatGPT should not stand in for a human case partner because it cannot read body language or adjust its questioning the way an interviewer does. We agree with most of that. Where we differ is on how much a well-instructed AI can add on scoring, which the later sections cover.
Our position is that ChatGPT helps most when you already know what a good answer looks like and you use it to generate volume, and helps least when you are asking it to tell you how good you are. The second job needs a standard, and a general chatbot has none unless you write one. The tabs below put the three main options side by side so you can see which job suits which tool.
Three ways to practise, and what each is for
Best for drilling and drafting. You can ask for any case in any industry, rerun a maths problem with new numbers, or paste in a structure and ask what is missing. It is free or cheap, always available, and endlessly flexible.
- Strong: maths drills, prompt generation, critique of a written structure, industry primers you then verify
- Weak: unprompted honesty, keeping case facts stable, scoring against a standard, hearing how you sound
- You supply: the case, the rubric, the strictness and the checking
Best for reps under speaking conditions. A purpose-built coach holds a fixed case, controls the pacing, interrupts, hands you exhibits and marks the transcript against a stated rubric. You give up flexibility: you run the cases it has, not the ones you invent.
- Strong: spoken delivery, answer order, pressure, consistent marking session to session
- Weak: a small library of cases, no read of body language, a marker that is still a language model and can be wrong
- It supplies: the case, the rubric and the pacing, and you supply the effort
Best for the things software has no standing to judge: whether you seem like someone a partner would put in front of a client, how you handle a person who is bored or sceptical, and the feel of a real audience. A peer is free and adds pressure. A former interviewer adds calibration.
- Strong: reading a room, a live person's style, stakes, and a sense of how you come across
- Weak: scheduling, patience with repetition, and feedback softened by friendship
- Scarce: a good session is hard to arrange, so spend it after the mechanics are solid
Most candidates end up using all three at different times, and the order counts for more than the ranking. The prep plan near the end of this guide sets out a sequence we would follow.
What is a general chatbot good for in case prep?
Maths first, because it is the cleanest win. Case interview maths is a set of mechanical skills (percentage changes, break-evens, margins, growth rates), and a chatbot will generate fifty variations in a minute and explain the steps when you get one wrong. You must check its arithmetic as well as your own, which turns out to be useful practice in itself.
Second, case prompts. Ask for a profitability case for a regional bakery chain, a market-entry question for a mid-sized insurer, a pricing question for a subscription product, and you get a workable opening prompt each time. Treat the prompt as raw material. What a chatbot produces after that opening is a different matter, and the failure section explains why we would not let it run the numbers unsupervised.
Third, critiques of a written structure, which is the use we recommend most. Write your issue tree in a few lines, paste it in, and ask the model to find overlaps, gaps and buckets that do not fit the question. Written structures are where a chatbot is on firmest ground, because the input is short, the task is bounded and you can judge the critique yourself against what MECE means. The value comes from the questions it raises, not the score it hands you.
Fourth, the sideline jobs that PrepLounge also lists: a quick industry primer before a case in an unfamiliar sector, a list of fit questions to answer out loud, and a second phrasing of a sentence you cannot get right. All of these are fine with one rule attached. Anything factual, especially any figure, gets checked against a source before you repeat it in a room.
What we would skip is asking a chatbot to decide whether you are ready. A model with no rubric and no memory of your earlier attempts has nothing to be right about, so its verdict tells you about the model's mood, not about you.
Where do chatbots fail as case interviewers?
Four failures come up again and again, in our own testing and in the candidate reports we have read. They are listed here in order of how much damage they do.
They let you off lightly
A poster in a PrepLounge thread on AI interview tools (posts run from April 2024 to January 2026, mostly 2024) put it bluntly: language models are not good at being critical, and whatever framework you show ChatGPT, it will praise it. Another said the advice was the generic sort you would get from a peer. That is one coach's opinion, but the research backs the mechanism. A 2023 paper led by Mrinank Sharma, titled 'Towards Understanding Sycophancy in Language Models', found that assistants tend to agree with the user's stated views and that this is driven in part by human feedback rewarding agreement. In April 2025, OpenAI rolled back a GPT-4o update in ChatGPT after users found it overly flattering, according to contemporary reporting of OpenAI's announcement. Newer models may do better, and you should not take our word or theirs: run the two-minute test described further down.
They drift off the case
A real case has fixed facts. The client sells 40 stores, the margin fell over three years, the exhibit shows one specific chart. A chatbot playing interviewer often changes those facts mid-conversation, answers a question you did not ask, or hands you the insight because you sounded uncertain. Fifteen minutes in, you are no longer solving the case you were given, and nobody is checking. A fixed case file with a hidden answer key is the boring fix, which is why purpose-built tools have them.
They are confident when wrong
Language models make arithmetic and recall slips and state them in the same tone as everything else. One reviewer of an AI interview tool in the same forum thread described its output as very short and often wrong. Another coach's line, 'garbage in, garbage out', appears in the ChatGPT thread. Whatever the exact rate on today's models, the risk is that you rehearse a wrong number with a calm voice telling you it is fine.
Spot the slip
A chatbot playing interviewer says: 'Revenue is 240 million and costs are 210 million, so the margin is 15%.' Is the interviewer right? Give the correct figure before you open the answer.
No. Profit is 240 minus 210, which is 30 million. Margin is profit divided by revenue, so 30 divided by 240, which is 0.125, or 12.5%.
15% matches neither reading: profit over costs would be 30 divided by 210, about 14.3%. The slip is small, but a candidate who accepts the figure carries it into every later step.
A strong candidate says the steps aloud, gets 12.5%, and says: 'I make that 12.5%, which is 30 on 240. Can you confirm the revenue figure?' Asking to confirm a number is never a deduction, and it is the behaviour you want to practise against a partner that gets things wrong.
This example is constructed to show the kind of slip, not a measured fault of any particular model.
They cannot hear how you sound
Typed practice hides hedging, pace, upward inflection and whether you led with the recommendation. Voice modes help, but a general chatbot's voice mode is built for conversation, not for marking delivery against a rubric, and it usually cannot tell you the pace at which you spoke or count your hedges. If delivery is your weakness, and for many candidates it is, a spoken tool with a transcript you can review is the better instrument.
Check yourself
You paste a structure into a chatbot and ask it to rate the structure out of 10. It says 8. What is the most useful next move?
That test is the most useful thing a candidate can do with any AI tool, ours included, and it takes about two minutes.
Which prompts make a chatbot a stricter case interviewer?
Prompts cannot remove a model's habit of agreeing with you, but they push it in the right direction, and the difference between a lazy prompt and a good one is large. Three that we would use follow, each aimed at one job. Copy them, change the bracketed parts, and rewrite them once you learn which instruction your chatbot ignores.
The first runs the interview. Its main features are a fixed fact sheet the model must not alter, one question at a time, and an instruction to wait while you think.
The second prompt is for scoring, and it does most of the job of making the marker strict. It supplies the rubric, demands quotations, and forbids the praise sandwich. Paste the anchors from our published rubric or your own, because a model with no anchors invents its own.
The third is a critic for written structures, which is the use we said chatbots are best at. It asks for problems, not a grade, which sidesteps the lenient-score habit.
Two cautions before you rely on these. First, test the marker prompt on a weak answer, as described above, and do not trust it until it produces a low score on demand. Second, a long conversation erodes instructions, so paste the fact sheet and rules again if the interviewer starts hinting. If you are curious how far a prompt can go, our comparison of prep tools sets a chatbot beside the dedicated products.
How do you judge the feedback from any AI case tool?
Judge the feedback, not the interface. A polished screen and a warm voice tell you nothing about whether the score means anything. The checklist below is the one we would use on a tool we did not build, and it applies to a chatbot with a strict prompt as much as to a product.
Ten checks for AI case feedback
0 of 10The third and fourth items are the ones we see tools fail most often, and they are worth a closer look. 'Be more structured' is a complaint about a category. 'When you said you would look at revenue and costs, you never said which you thought was driving the drop, so the structure had no hypothesis' points at a moment you can go back to and change. Case interview feedback goes through that difference at length, and it applies whoever or whatever is giving the feedback.
Check yourself
A tool returns four feedback lines on a market-sizing case. Which one would you trust and act on?
The first item on the checklist, a published rubric, deserves the most attention because everything else depends on it. Without written anchors, a score is a mood. With them, you can read the 1, 3 and 5 for a skill, find where you sit and see what the next level looks like. Here is the anchor for framing from the rubric our coaches mark against.
How this is marked · Analytical thinking
Framing
Whether the candidate builds a structure that is MECE, hypothesis-led, and adaptive
- 1
Weak
Reaches for a memorised framework with no fit to the problem, or buckets with no logic.
- 3
Sound
A sound, broadly MECE structure that fits the problem and is usable, even if a layer is thin or the hypothesis is only implied.
- 5
Outstanding
A structure tailored to this problem (any valid shape counts), genuinely MECE drivers, a stated or clearly-implied hypothesis, and adapts when new information lands. Judge the structure the candidate had committed to by the time they finished building it, do not penalise missing depth on a first layer they were still in the middle of laying out, or for asking to verify direction before going deeper.
Scores run 1 to 5 per skill. The first-round bar is an average of 3.5, so a 3 is sound but not yet enough on its own.
Anchors like these are what let a marker say 3 instead of 'good structure', and let you check the marker
See all 14 skills in the published rubricNotice what the five-point anchor asks for: a structure tailored to this problem, a hypothesis, and adaptation when new information arrives. Notice also the instruction to judge the structure a candidate had committed to by the time they finished building it, so a first layer still being laid out is not marked as thin. Guardrails like that stop a marker from penalising a candidate for something the interview would not penalise. A rubric is a set of rules for marking, and the rules include what not to deduct for.
The second skill worth reading is rigor, because it holds the separation between a wrong number and a missing sense-check that a lazy marker collapses into 'work on your maths'.
How this is marked · Analytical thinking
Rigor
Whether the candidate's quantitative work is precise, traceable, and sense-checked
- 1
Weak
Avoids the math, or is wrong by an order of magnitude with no recovery.
- 3
Sound
Gets the math right (at least order-of-magnitude) but slowly, or only sense-checks when prompted.
- 5
Outstanding
Correct on the first pass, states the calculation aloud, and sense-checks the result against a known anchor on their own initiative.
Scores run 1 to 5 per skill. The first-round bar is an average of 3.5, so a 3 is sound but not yet enough on its own.
Note the last sentence of the 5 anchor: needing a number repeated is never a deduction
See all 14 skills in the published rubricTwo candidates with the same wrong answer can score differently here, because one caught the slip on their own initiative and the other never checked. A tool that reports only 'incorrect calculation' has thrown away the more useful half of what happened.
Why is AI feedback so forgiving, and how does our marking handle it?
Two causes stack on top of each other. The first is the sycophancy tendency described earlier, where a model trained to be helpful and liked drifts toward agreement. The second is design: many tools grade against an unstated standard, or worse, against an implied average of whoever has used the tool, and an average of self-selected practisers gives a comfortable middle score to nearly everyone. A candidate who scores 4 out of 5 every week has learned that the tool is easy to please, not that they are ready.
The way out is to fix the standard in advance and mark against it. Our scoring is criterion-referenced, meaning each skill is read against written anchors and against a fixed first-round bar of 3.5 out of 5, not against a curve and not against other candidates. Nobody's score depends on how anyone else did that week. The overall figure is a weighted average across three dimensions (analytical thinking at 50 per cent, presence and communication at 25, and track record and drive at 25 in the full interview), and the 30-minute short case, which has no behavioural part, drops the last dimension and reweights the other two.
That design is what allows low scores. In our own validation runs on real transcripts, weak cases came out between roughly 1.7 and 2.7 out of 5, and a constructed excellent transcript reached about 4.6, so the scale is used at both ends. Those are checks we ran on our own marker with our own transcripts, not an independent audit, and they show the ceiling and floor exist rather than that any single score is exact. A language model is still doing the marking, it can misjudge a case, and we would rather you tell us when a score looks wrong than assume it is right.
A fixed bar also carries guardrails against being harsh for the wrong reasons. Asking the interviewer to repeat or confirm a number is never a deduction, and on the listening skill it counts as a strength. The framing skill is judged on the structure the candidate had committed to by the end of the build, not on a half-finished first layer. And a case where the candidate barely got going does not receive a low score at all: the marker declines to score it, because a fabricated 1.2 on a case that never started is another way of being unhelpful. The hire labels attached to the score are developmental, so there is no verdict of 'reject', only a read of how far a skill sits from the bar.
What can AI practice not replace?
Three things, and we would not pretend otherwise. The first is reading a room. A human interviewer's mood changes within a case: they lean in when you say something sharp, fall silent when you lose them, ask a question because they sensed hesitation. You learn to notice and adjust to that only with a person opposite you, and a voice model, however patient, is not showing you a face.
The second is a particular partner's style. Interviewers vary in how much they interrupt, whether they run the case themselves or wait for you to lead, and how much they care about elegance versus speed. Interviewer-led versus candidate-led cases covers the difference, and firm-specific habits are covered in the guides on McKinsey, BCG and Bain. An AI coach can simulate either format, and it cannot tell you how the person you will meet on the day prefers to run theirs.
The third is stakes. Nothing you do with software carries the weight of an actual round, where a bad opening cannot be rerun and someone will remember it. Some candidates freeze in the real thing after acing every practice case, and the fix is exposure to a person whose opinion counts, whether that is a friend from your programme or a former interviewer you pay. Both are worth arranging before the round, and the final-round guide covers what changes at the end of the process.
PrepLounge's July 2025 article makes the same point about human feedback, that AI cannot fully replicate the reasoning and experience of an MBB interviewer. We hold that view about our own product. AI practice is the cheapest way to get through the early volume, and it is a poor substitute for a person at the end.
How should AI practice fit into a two-week prep plan?
Sequence counts for more than any one tool. The plan below assumes two weeks, one hour a day and no partner guaranteed on demand, and it moves from cheap and flexible to scarce and calibrated. It is our suggested shape, not something we have tested against outcomes, so adjust it to your dates. If your interview is closer than two weeks, cut the drills and keep the human session.
A fortnight that uses each tool for what it is good at
- Days 1 to 3
Drill maths with a chatbot
Ten minutes a day of generated maths problems. Do each one aloud, then check the chatbot's working as well as your own, and note every time you sense-checked without being asked.
- Days 3 to 5
Critique written structures
Write six structures for six generated prompts and use the third prompt above to find holes. Compare its critique with the issue tree guide and overrule the chatbot where you disagree.
- Days 5 to 9
Spoken cases with a marked report
Run a spoken case each day with a coach that has a fixed case and a stated rubric. Read the evidence lines, pick one skill to change, and let the next case test that change.
- Days 10 to 11
Book a human mock
A peer or a former interviewer. Tell them nothing about your AI scores, and ask for the moment they would have interrupted you.
- Days 12 to 14
Rerun your weakest skill
One or two more spoken cases aimed at whatever the human mock and the scorecards agree on. Stop early on the last day and sleep.
One habit separates the candidates who improve from those who do not: they read the feedback before starting the next case. Ten cases run back to back without changing anything is repetition, and repetition alone tends to keep a score where it is. Case interview mistakes lists the recurring ones worth watching for, and practising alone covers the routine around them.
For full disclosure, since we make one of these: MBB Ready runs spoken cases with AI coach personas, marked against the published rubric described above, and the first case is free. There is also a short taster without signing up. It cannot tell you how a particular partner will read you, and nothing on this page changes that, including the product we built.
Common questions
Can ChatGPT run a full case interview?
It can run one if you give it a strict prompt with a fixed fact sheet, and both ChatGPT and Claude now have voice modes. The weak points are drifting off the case facts, agreeing with weak answers and making arithmetic slips, so you supply the standard and check the numbers yourself. A test with a deliberately bad answer shows quickly how far you can trust it.
Which ChatGPT prompt is best for case interview practice?
Use three prompts for three jobs. One runs the case with a hidden fact sheet, one question at a time and no hints. One marks a transcript against anchors you paste in, with quoted evidence. One finds holes in a written structure without rating it. This guide contains all three, ready to copy. Test the marker prompt on a bad answer before you trust its scores.
Why is AI case interview feedback so generic and forgiving?
Language models trained to be helpful tend to agree with the user, which researchers call sycophancy, and many tools mark against an unstated standard, which produces comfortable middle scores. The fix is a written rubric, a fixed bar, and feedback that quotes what you said. If a tool still praises a deliberately weak answer, treat its scores as noise.
How do I know if an AI interview tool's feedback is reliable?
Check whether it marks against a published rubric, quotes your words, names a skill and a moment, separates wrong numbers from missing sense-checks, and scores a weak answer low. Then run the same transcript twice and see whether the score is stable. A tool that fails the weak-answer test is not marking, however good the interface looks.
Is it better to practise with AI or with a human partner?
They do different jobs. AI is better for volume, maths drills, written structure critiques and consistent marking at any hour. A human is better for reading a room, adapting to a specific style and feeling real stakes. Use AI early for repetitions and keep at least one human mock, ideally with a former interviewer, before a real round.
Can AI tools replace a case interview coach?
Not for everything. AI can drill and mark consistently, but a good human coach reads how you come across, notices patterns across many candidates and calibrates you against a real hiring team. PrepLounge's own guidance makes the same point. Treat AI as the cheap volume layer and a human as the calibration layer.
Does asking the interviewer to repeat a number hurt my score?
It should not, and in our rubric it never does. Asking an interviewer to repeat or confirm a number you did not hear is normal, and on the listening skill it counts as a strength. The marker looks at your reasoning, your arithmetic and whether you sense-check, not at whether you needed a figure repeated.
Sources
- PrepLounge, forum thread: using ChatGPT and AI for case practice (coach and candidate posts, February to March 2025)
- PrepLounge, article: AI for interview preparation (published July 2025)
- PrepLounge, forum thread: AI interview tools for MBB (candidate and coach reports, April 2024 to January 2026)
- Sharma et al., Towards Understanding Sycophancy in Language Models (arXiv, 2023, revised 2025)
- MacStories, report on OpenAI's April 2025 GPT-4o sycophancy rollback (secondary reporting of OpenAI's announcement)