Digital transformation and AI
Digital transformation and AI

Ask a frontier AI model to critique a market entry decision, explain a valuation for a multinational firm facing currency shock pressures, or advise a hospital executive on a fraught promotion decision, and odds are it will get most of the answer right. A new benchmark built by researchers at the University of Pennsylvania, Carnegie Mellon University, and Harvard, used 238 real MBA case studies spanning eighteen business disciplines. This benchmark, graded against the same rubrics instructors use to grade their own students, found that today's leading models score above 87% on the kind of open-ended analytical reasoning business schools have spent a century teaching and perfecting.

I spoke to two of the authors of this benchmark, Kartik Hosanagar of the Wharton School of Business and Ramayya Krishnan of the Heinz College at Carnegie Mellon University, about the implications of their work and why developing such a benchmark is important.

Hosanagar and Krishnan explained that for a lot of entry level knowledge workers in organizations, work they get assigned is typically along the lines of "develop a plan for market entry and name acquisition targets by xx(deadline)". Think of an entry level strategy analyst in a consulting firm. Knowledge work is specified in ambiguous terms, without a defined formula and without specifying a right answer. Existing AI benchmarks do not test this aspect of knowledge work at all. Hosanagar and Krishnan highlighted the precedent from medicine where medical lexicon and knowledge encoded in large language models (LLMs) can be better evaluated on case challenges from the New England Journal of Medicine (NEJM) instead of benchmarking on the United States Medical Licensing Examination (USMLE) which typically contain multiple-choice questions.

The current study by Hosanagar, Krishnan, Callison-Burch and Lakhani extends such an approach to business school cases. This provides a more realistic benchmark to assess how AI can perform in a real-world entry level knowledge work setting that requires active, diagnostic reasoning than static understanding. Their approach looks at 238 licensed business cases with 615 questions across 18 disciplines, each graded against the instructor's own reference solution. They simulated fictional vs. real firms, numerical vs. non-numerical data, subjective vs. objective criteria, and mapped to occupational taxonomies of skills and professions from O*NET, and their approach is robust to contamination across five major pretraining corpora (C4, Pile, RedPajama, Dolma, DCLM) to ensure that the benchmark data do not inadvertently leak into the text used to train the LLMs.

Two years ago, the same benchmark would have stumped AI models: one model's score jumped 23 percentage points since 2024. The implication isn't that AI can replace a strategist or a CFO tomorrow. It's that the skill that business education has always prized most - that of producing a defensible, structured analysis under uncertainty - is now something a machine can draft in seconds. The open question is what that does to how we train talent, and who gets hired to do the thinking work that's left.

That headline number, though, hides a more interesting story. The researchers also scored responses on a stricter standard: not just partial credit for hitting some of the rubric, but full credit only when a model nailed every single expected element of an instructor's model answer.

Hosanagar and Krishnan emphasized a key aspect of real-world problem solving is that we don't often have a perfect solution; rather, what the strategy analyst in the consulting company is working on is to find an acceptable solution given the constraints of time and resources. They incorporated this aspect of decision making into their benchmark by considering the difference between partial-credit scoring vs. full credit (complete answer) scoring. The top AI models score 87%+ under partial-credit ("Standard") scoring with a 6.8-point spread between the best and worst frontier model. But under stricter "Complete Answer" scoring (every rubric item must be hit), even the best model completes fewer than half of questions fully.

On that measure, even the best-performing model completed fewer than half the questions in full. The gap between "impressively competent draft" and "complete, decision-ready answer" turns out to be exactly where the frontier still lives. And the frontier is not evenly distributed: in their evaluation, models are already performing great at structured, bounded tasks like financial calculations, but stumble on open-ended advisory judgment. In other words, the "what should we actually do here" moments that don't come with a formula are difficult for frontier models. That distinction matters more for recruiting and business education than the aggregate score does.

As Krishnan and Hosanager explain this "AI writes exceptional first drafts, not yet gold-standard final drafts." The discipline matters more than question type: the gap between easiest and hardest discipline (80.1% to 95.0%) exceeds the gap between models. Open-ended advisory work (e.g., "identify business opportunities," "advise on financial matters") is the hardest, while structured/quantitative tasks are near ceiling. Fewer than 7% of questions defeat every frontier model; a "best-of-three" oracle reaches 92.6% and they find that different models miss different things.

Where the frontier actually sits shows up in the discipline-level variation (Business & Government Relations are at near-ceiling; Marketing and Operations weaker); the O*NET occupational mapping shows "open-ended advisory" work (such as identifying opportunities and advising stakeholders) as the hardest activities, versus structured analysis as near-solved.

The most interesting thing here is not just the pace of change. Not only we see a 23-point generational jump in two years, but it's also broad-based (not just numerical reasoning, which is the "easy" capability everyone expected to improve first). The implications for business education are clear; there is a shift from "can you produce an analysis" to "can you audit, stress-test, and improve one", which calls into question the future of the case-study method of teaching. Hosanagar and Krishnan explain this shift in terms of the change in skills required upstream as well as downstream. "If AI is changing the process of production (how we solve a problem), then how to we evaluate what is the problem to be solved?"

Equally interesting are the implications for recruiting and entry-level roles. Typically, junior analyst work is seen as the traditional training ground for developing judgment and reasoning. This benchmark demonstrates the case for new apprenticeship models that build verification and supervision skills earlier.

For educators, studies such as this point to the fact that scarce skill shifts from producing analysis to verifying it; teaching burden increases rather than decreases. Hosanagar mentioned an anecdote whereby one "needed years of experience building different types of valuation models before you could pick up someone else's model and assess it". For organizations, this is a pointer that value migrates upstream (framing the right question) and downstream (building consensus, owning the decision) as the middle "draft" gets commoditized. This raises the need for a "new apprenticeship problem" that replace earlier models of training where junior roles historically built judgment through years of learning by doing.

The skills that used to be earned through years of producing first drafts such as spotting a weak assumption, sensing when a recommendation is half-baked now need to be taught directly, and earlier, because the production work that used to teach them is rapidly changing. Time for a new version of the Apprentice: AI Edition?

This article was originally published on Forbes.com