AI systems began scoring at professional levels on demanding exams. Large language models passed medical, legal, and business tests once used to certify humans. The results sparked debate over capability and meaning.
Professional-level scores
The bar was high. Models passed licensing-style exams. Performance surprised many.
Broad domains
Range is wide. Medicine, law, and business tests were cleared. Breadth impressed.
Rapid gains
Speed startled. Scores jumped across model generations. Progress was fast.
Not understanding
Caution is loud. Passing exams is not true competence. Limits persist.
Real-world gap
Nuance matters. Exams differ from practice. Deployment needs care.
A capability debate
Discussion followed. What the scores mean is contested. The field reflects.
The bottom line
Large language models passed professional medical, legal, and business exams, with scores rising fast across generations. The results are striking but do not equal real competence. Their meaning is debated.