Can AI estimate calories from a photo?
Yes, as an estimate, and the weak step is not recognising the food. A vision model names what is on the plate, guesses how many grams of each item there are, then converts those grams to calories. The gram guess decides the result. When we ran the model inside our own app on 10 cafeteria plates whose ingredients had been weighed, 3 came within 25% of the true calories, and the misses ran from 52% under to 139% over.
Below are the three steps behind every photo estimate, our test plate by plate, what independent studies found for ChatGPT, Claude, Gemini and four commercial apps, and the extra information that cut one model's error by more than half.
Disclosure
We publish Calso, a photo calorie app, and the test on this page is of our own app. The result does not flatter it. The plates, photos and weights come from a public dataset, linked at the foot of the page, so every row of our table can be checked against it.
How does AI estimate calories from a photo?
In three steps, and only the first is what most people picture. A photograph contains no weight, so the model has to supply one before any nutrition data can be used.
| Step | What the model does | Where it goes wrong |
|---|---|---|
| 1. Name the components | Lists each food it can see, including sauces, dressings and oil | Items hidden under others, fat it cannot see |
| 2. Estimate the grams | Scales each item against objects of known size, such as the plate and cutlery | A flat photo has no depth, so height and density are guessed |
| 3. Convert to calories | Multiplies each mass by per-100 g values for the food as it was cooked | Frying adds oil the photo does not show |
Here is step two in practice. Calso's instructions tell the model that a dinner plate is 26 to 28 cm across and a fork 19 cm, and require a mass in grams for every component before any calorie figure. The app then adds up the components itself instead of asking the model for a total. The model is OpenAI's, which our Privacy Policy names as the recipient of meal photos.
Research systems attack step two with hardware. Nutrition5k, a dataset Google Research released with a 2021 paper, scanned 5,006 cafeteria plates on a custom rig that captured an overhead depth image alongside the colour photo where available, and weighed every ingredient. A single ordinary photo carries no depth measurement at all.
Which step does AI get wrong most often?
The grams. Our first test made that plain, because it separated the arithmetic from the portion.
Our first instructions asked the model for calories directly. On 30 weighed plates, 8 landed within 25% of the true figure, 24 came back too high, and the median estimate was 60.5% over. Yet every one of the 30 breakdowns was internally consistent: protein, carbohydrate and fat added up to within 10% of the model's own calorie figure. The sums were right. The quantities were not. One plate holding spaghetti, a slice of pizza, small potatoes and a little caesar salad weighed in at 398 kcal and came back as 1,056 kcal.
The worst cases showed the model recognising each dish and pricing a standard serving of it, instead of reading how much was on this plate. Rewritten to force grams first, the instructions brought the median on the first 10 plates of the same set to 24% over. Only 3 of those 10 landed within 25%.
How accurate was our app on weighed meals?
3 of 10 plates came within 25%, which fails the bar we set ourselves of 8 in 10. Each row is one plate from Nutrition5k, named by its three heaviest ingredients, photographed from above and read once by the model and instructions Calso uses.
| Plate | Weighed | Estimated | Calories off | Grams off |
|---|---|---|---|---|
| Spinach, eggplant, tofu | 400 kcal, 442 g | 194 kcal, 330 g | -52% | -25% |
| Mixed greens, avocado, parmesan | 250 kcal, 153 g | 136 kcal, 108 g | -46% | -29% |
| Salmon, cherry tomatoes, lentils | 238 kcal, 161 g | 237 kcal, 167 g | 0% | +4% |
| Chicken thighs, wheat berries, broccoli | 483 kcal, 409 g | 517 kcal, 402 g | +7% | -2% |
| Chicken, broccoli, spinach | 286 kcal, 280 g | 306 kcal, 285 g | +7% | +2% |
| Cauliflower, cheese pizza, roast potatoes | 540 kcal, 336 g | 765 kcal, 425 g | +42% | +26% |
| Chicken, cucumbers, pork | 285 kcal, 209 g | 463 kcal, 320 g | +62% | +53% |
| Berries, pizza, beef | 311 kcal, 196 g | 673 kcal, 350 g | +116% | +79% |
| Waffles, tofu, sweet potato | 544 kcal, 361 g | 1,227 kcal, 570 g | +126% | +58% |
| Caesar salad, roast potatoes, pizza | 398 kcal, 350 g | 950 kcal, 520 g | +139% | +49% |
Four things stand out.
- Calories follow grams. The three plates within 25% were also the three whose total weight the model got within 4%. Every large calorie miss sits on a weight miss of 25% or more.
- Dense plates come back high. The waffles plate weighed 361 g and read as 570 g, so 544 kcal became 1,227. The plate that read 1,056 kcal under the first instructions still reads 950.
- Oily salads come back low. The spinach, eggplant and tofu plate held 33 g of fat. The model found 10 g and put the dressing at 5 g. The greens, avocado and parmesan plate came back with no avocado listed at all.
- The same photo does not get the same answer. That spinach plate, 400 kcal on the scale, has read 326, 245 and 194 kcal across three calls with near-identical instructions. Rewording cannot remove a spread that wide.
How we ran it. Plates were taken from Nutrition5k in a fixed order set by a hash of each dish ID, so none could be dropped after seeing the result, and limited to meal-sized plates: 150 to 1,000 kcal, at least 100 g and at least two ingredients. The model was OpenAI's gpt-5.6-luna, run on 12 September 2026. Ten plates is a screening, not a verdict. We stopped at ten because even a perfect score on the next 20 could not lift the set to 24 of 30.
Are ChatGPT, Claude and Gemini better at it?
Not in the published tests, and commercial apps show the same weakness. Two independent evaluations against weighed food are worth knowing.
| Study | What was tested | Calorie result |
|---|---|---|
| University of Gothenburg, 2025 | ChatGPT-4o, Claude 3.5 Sonnet and Gemini 1.5 Pro on 52 standardized food photos, in three portion sizes | ChatGPT and Claude: 35.8% average error. Gemini: 64.2% to 109.9% across nutrients |
| NIH (NIDDK), presented July 2026 | MyFitnessPal, Lose It!, Cal AI and Appediet on 102 meals from a metabolic kitchen, ingredients weighed to 0.1 g | All four underestimated by about 250 to 345 kcal on average, and fat by about 30 g |
The Gothenburg team found all three models underestimated, more so as portions grew, and concluded they are not yet suitable where precise quantities matter. The NIH team, who presented at the American Society for Nutrition's NUTRITION 2026 meeting, found fat was the most underestimated, and early results on more than 200 further meals suggest high-fat ketogenic meals are the hardest. Their advice, from NIDDK researcher Aaron Hengist: people using a photo app "without adjusting the portions or entering the amounts of food should take the results with a grain of salt."
Those results match ours on the mechanism, if not the direction. Our small cafeteria plates came back too big, their larger portions came back too small, and both fit an estimate pulled toward a typical serving. Fat is where our low readings and theirs overlap. Results for those model versions do not automatically carry over to newer ones, which is why we test our own.
What makes an AI calorie estimate more accurate?
Telling it what the photo cannot show. A 2025 study in Nutrients gave ChatGPT-5 the same 195 dishes four times, with a different amount of context each time.
| What the model received | Average error, calories |
|---|---|
| Photo only | 30.51% |
| Photo, plus type and amount of fat and sweetener, dairy fat content and type of meat | 24.35% |
| Photo, plus a full ingredient list with amounts | 13.92% |
| Ingredient list with amounts, no photo | 18.05% |
Two lessons. Words about fat and quantity cut the error, and the photo still earns its place: taking it away made the estimate worse. Part of that study's reference values came from recipe and database figures rather than a scale, so compare these rows with each other rather than with the weighed-food studies above.
Whichever app you use, that turns into four habits:
- Say the fat. If your app accepts a note, name the oil, butter or dressing and roughly how much. It is the context a photo hides best.
- Correct the weight of the densest item. In our test the calorie misses tracked the gram misses, so fixing the grams of cheese, pizza or meat does more than nudging the total.
- Keep the plate and cutlery in frame. Our instructions and the Gothenburg study's prompt both use them as the ruler. Crop them out and the model has nothing to scale from.
- Weigh what is dense, photograph the rest. Foods above roughly 350 kcal per 100 g are where a misjudged portion costs the most, which is the subject of counting calories without a food scale.
In Calso, the correction happens on the meal screen: every scanned meal opens to its calories, protein, carbohydrate and fat, and each figure can be edited. Calso does not take a written note with the photo, so the fat you know about goes in as a correction afterwards.
FAQ
Can ChatGPT count calories from a picture?
It will give you a number. In a 2025 test against weighed food, ChatGPT-4o missed the calories by 35.8% on average. A separate 2025 study found ChatGPT-5 averaged 30.51% error from the photo alone and 13.92% when it was also given the ingredients with amounts. The more you tell it, the closer it gets.
Why does the same photo give different calories each time?
Because the model makes a fresh judgement about portion size on every call, and it does not make the same one twice. In our test one plate of 400 kcal read 326, 245 and 194 kcal across three calls. A rescan is a new guess, not a check, so correct the number instead.
Do AI calorie counters overestimate or underestimate?
Both, depending on the plate. The NIH test of four apps found underestimation of about 250 to 345 kcal per meal on average, and the Gothenburg study found underestimation that grew with portion size. In our own test, small energy-dense plates came back high and oily salads came back low.
Is describing a meal in words better than a photo?
Words with amounts beat a photo alone: 18.05% error against 30.51% for ChatGPT-5 in the 2025 Nutrients study. Both together did best, at 13.92%. If you cannot list every ingredient, a photo plus a few facts it cannot see, such as the type and amount of fat, still brought the error down to 24.35%.
Which foods do AI calorie counters get most wrong?
Anything whose calories are not visible. Fat is the common thread: the NIH researchers found fat underestimated by about 30 g on average and ketogenic meals hardest, and our worst low readings were salads carrying oil, cheese or avocado. Mixed plates are also harder than single foods, as the accuracy research shows.
Three things to take away
- AI can estimate calories from a photo, but it is guessing the weight, and the weight decides the answer.
- Against weighed food, published errors for chatbots and apps sit around a third, and our own app got 3 of 10 plates within 25%.
- Tell it what it cannot see, the fat and the amounts, and correct the grams of dense food rather than rescanning.
Calso, with numbers you can correct
Photograph your plate, get calories and macros logged against targets built from your body and your goal, and change any figure on the meal screen when you know better. The home page lists what the app does, and what it does not.
Sources
- Nutrition5k dataset, Google Research, CC BY 4.0, released with Nutrition5k: Towards Automatic Nutritional Understanding of Generic Food, CVPR 2021. The 10 plates in our table are dishes 1560368815, 1566849269, 1564773847, 1560974009, 1566413419, 1565036922, 1561664042, 1562703408, 1565021220 and 1562963676, in table order.
- Performance Evaluation of 3 Large Language Models for Nutritional Content Estimation from Food Images, Fridolfsson et al., Current Developments in Nutrition, 2025.
- Image-Based Dietary Energy and Macronutrients Estimation with ChatGPT-5, Rodríguez-Jiménez et al., Nutrients, 2025, 195 dishes.
- Photo-based calorie-tracking apps fall short in estimating meal nutrition, American Society for Nutrition, via News-Medical, 26 July 2026.