Millions of people now point a phone camera at a plate of food and let an AI model estimate the calories, but new research presented at NUTRITION 2026, the American Society for Nutrition’s annual meeting held July 25-28, 2026 in National Harbor, Maryland, suggests those estimates are frequently and systematically too low. The study, led by Aaron Hengist, a postdoctoral visiting fellow, and Olivia Charles, a postbaccalaureate research training fellow, both at the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK), part of the National Institutes of Health, found that four widely used photo-based nutrition apps underestimated calories by an average of 250 to 345 per meal.
How the test was built
Rather than relying on self-reported meals or restaurant menu data, the NIDDK team built its test around a controlled metabolic kitchen, where ingredients for 102 meals were measured to a precision of 0.1 gram before being photographed and fed into four apps: MyFitnessPal, LoseIt!, CalAI, and Appediet. Because the true calorie and macronutrient content of each meal was known down to the gram, researchers could measure exactly how far off each app’s AI-generated estimate landed rather than comparing one estimate against another. After the initial 102-meal test, the team expanded the analysis with more than 200 additional meals to probe which factors most affected accuracy, such as mixed dishes, unusual plating, or portion size.
Where the errors concentrated
The most consistent failure point was fat content, which the four apps undercounted by roughly 30 grams on average, a discrepancy large enough to meaningfully change a meal’s true calorie load given that fat carries nine calories per gram versus four for protein or carbohydrate. Carbohydrate estimates, by contrast, were comparatively more consistent across the apps, suggesting the computer-vision models are better at recognizing and sizing starches and sugars than at judging how much oil, butter, or hidden fat is present in a mixed dish, information that is often invisible to a camera lens looking down at a plate.
The researchers’ own caution
Hengist was notably measured in interpreting his own results, telling reporters that users “should take the results with a grain of salt,” while adding that the apps “tend to underestimate calories, especially from fats.” That caveat matters for how the finding should be read: the research was presented at a scientific conference rather than published in a peer-reviewed journal, and the authors themselves describe the results as preliminary rather than final, meaning the precise numbers could shift once the work goes through formal peer review.
A counterpoint from the industry building better scanners
The NIDDK findings landed just as newer entrants were pushing in the opposite direction. On August 12, 2026, Orange County-based startup Nourish Lens launched an updated AI nutrition scanner aimed specifically at the portion-estimation and mixed-meal problems the NIH study highlighted, part of a wider wave of apps, including Yuka, which scans more than 1.5 million food and cosmetic products by barcode using a scoring system weighted 60% toward nutritional quality, 30% toward additive presence, and 10% toward organic status, betting that better computer vision and larger training datasets can close the accuracy gap the NIH researchers documented. Industry data suggests the challenge is real: separate estimates indicate roughly 82% of AI nutrition app initiatives fail to reach clinical-grade accuracy and durable user adoption, citing insufficient training data, poor portion-size estimation, and difficulty handling mixed or layered meals, the exact weaknesses the NIDDK metabolic-kitchen test was designed to expose.
Why the gap matters beyond a few missed calories
For a casual user tracking weight loosely, an average underestimate in the 250-to-345-calorie range per meal is unlikely to be noticed day to day. But for patients managing diabetes, cardiovascular risk, or working with a clinician on a structured nutrition plan, where an app’s output feeds directly into medical decisions, a systematic undercount compounded across three meals a day could mean a person believes they are consuming meaningfully fewer calories and less fat than they actually are, undermining the exact clinical goal the technology is meant to support. That risk is precisely why NIDDK, a research arm focused on diabetes and metabolic disease, chose to formally test consumer apps rather than leave the accuracy question to app-store reviews and marketing claims.
What happens next
The NIDDK team has signaled it intends to pursue formal peer-reviewed publication of the expanded 200-plus meal dataset, which should clarify whether the errors are consistent across cuisines, portion sizes, and plating styles, or concentrated in specific categories like fried foods and sauces. Meanwhile, app makers face a choice: absorb the criticism and treat it as a roadmap for model improvement, as competitors like Nourish Lens appear to be doing, or contest the methodology. Either way, the episode is a reminder that AI tools embedded in daily wellness routines are only as trustworthy as the ground-truth data used to check them, and that a controlled metabolic kitchen, however unglamorous, remains one of the few ways to actually verify what a camera-based AI model is getting wrong.