Millions of people now ask ChatGPT, Claude, or Gemini what to eat this week. The plan shows up in seconds, formatted like a coach wrote it, complete with grams, macros, and a neat daily total. The plan is easy to get and easy enough to trust as a starting point, so people keep the scaffold and adjust the parts that feel off.

Gemini building a 1,500-calorie day, each ingredient in grams.
It is also why the old failure mode still matters. Early AI meal plans were famous for confident nonsense, from missing allergens to phantom micronutrients to calorie totals that fell apart under a second pass. The products improved. The open question is how much, and whether the three assistants people actually use fail in the same ways.
So we tested the consumer apps directly. In an August 2026 benchmark across ChatGPT, Claude, and Gemini, calorie math held for two of the three. GPT-5.5 Instant and Claude Sonnet 5 each landed within 5 percent of the requested calorie target in all 12 of their targeted runs. Gemini 3.6 Flash hit that same 5 percent bar in 5 of its 12 runs, and its misses were not small. One 2,500-calorie plan came back at an independently recalculated 3,216.6 calories.
Calorie accuracy is only one part of a useful meal plan. Our guide to the best AI nutrition apps covers the wider product landscape. How Accurate Is AI Food Logging? covers the reverse problem, estimating what you already ate. This piece covers what our tests of the current apps found, what peer-reviewed studies have found since 2023, and what still needs a manual check before you build a week of eating around an AI-generated plan.
01We tested the assistant apps most people use
Most AI meal-plan research evaluates a model through an API or a study-specific interface. We tested the consumer assistants people open every day: ChatGPT with GPT-5.5 Instant, Claude with Sonnet 5, and Gemini with Gemini 3.6 Flash.
These assistants include product instructions, safety policies, interface defaults, and other behavior layered around the underlying model. Our benchmark therefore measures the answers people receive from the finished assistant apps.

Claude refused the 1,000-calorie plan all three times.
On August 4 and 5, 2026, we gave each assistant the same seven prompts three times in fresh temporary or incognito chats. The 63 responses tested whether the assistants could hit calorie targets, break compound recipes into measurable ingredients, exclude allergens, handle an aggressive weight-loss request, and recognize the nutrient limits of a vegan plan without supplements or fortified foods.
The sections below compare those results with meal-planning studies published between 2023 and 2026. The complete testing and calorie-recalculation method appears later for readers who want to inspect or repeat the benchmark.
02ChatGPT and Claude hit calorie targets Gemini missed 7 of 12
ChatGPT and Claude stayed close to the target every time. Gemini's answers changed from run to run.
| Product | Targeted runs | Median absolute target error | Mean absolute target error | Runs within 5% | Largest miss |
|---|---|---|---|---|---|
| ChatGPT GPT-5.5 Instant | 12 | 2.50% | 2.37% | 12 of 12 | 4.27% |
| Claude Sonnet 5 | 12 | 1.37% | 1.30% | 12 of 12 | 2.65% |
| Gemini 3.6 Flash | 12 | 6.42% | 7.82% | 5 of 12 | 28.66% |
GPT-5.5 Instant and Sonnet 5 never missed their target by more than 4.27 percent across a 1,500-, 2,000-, or 2,500-calorie daily plan or a 500-calorie stir-fry recipe. Gemini's seven misses included a 2,500-calorie plan that added up to 3,216.6 calories and a 500-calorie stir-fry that added up to 574. That is not a rounding error. Follow that plan as written and the day lands closer to a bulk day than a 2,500-calorie day. One Gemini answer claimed 2,004 calories. Its own protein, carb, and fat numbers added up to about 2,200.
03Peer-reviewed AI meal plan accuracy findings 2023 to 2026
Much has changed since 2023. Consumer assistants behave differently from earlier models and have largely improved with each model upgrade. Read together, the studies show which failures earlier research exposed and how the 2026 apps handled the same questions.
2023. Niszczota and Rybicka built 56 elimination diets with ChatGPT for hypothetical people with specific food allergies and checked every meal for forbidden ingredients.1 Four of 56 diets contained the forbidden allergen, most often almond milk in a nut-free diet. Among the 14 intentionally 1,000-calorie diets in that study, only one included a warning about the calorie restriction.
2024. Hieronimus and colleagues evaluated 108 one-day meal plans generated by ChatGPT 3.5 and Bard across omnivorous, vegetarian, and vegan patterns.2 Vitamin D and fluoride fell below the dietary reference intake in every plan tested. Vitamin B12 was insufficient only in the vegan plans, and ChatGPT recommended a B12 supplement in 5 of 18 vegan-plan instances while Bard never did.
2025. Three studies moved from allergy and micronutrient failures toward the calorie question directly. Aslan and Sozlu found ChatGPT 3.5 under-shot 1,500- and 2,000-calorie daily plans, while 2,500-calorie plans landed near target.3 In a weight-loss plan comparison across Gemini, Microsoft Copilot, and ChatGPT 4.0, Gemini was the least consistent chatbot. Five of ten plans missed the stated target by at least 20 percent. ChatGPT 4.0 stayed closest.4 Macronutrient and fatty-acid balance scored weakest across all three chatbots. On compound-ingredient decomposition across 15 meal plans and 99 ingredients, correct weight matching ranged from 55 percent with Mixtral 8x7B to 87 percent with Llama-3 70B.5
2026. Harder prompts still broke down. Dobrowolski's 3(LM)Diet study asked ChatGPT (GPT-5.5), Gemini (Gemini 2.5 Flash), and Microsoft 365 Copilot for 35-day plans at 2,000, 2,500, and 3,000 calories under free-account, non-expert conditions in Polish.6 All three under-shot energy by 284 to 546 kcal depending on the model and target, and none produced a single 2,000-calorie plan that hit its exact target. Gemini's plans were the most deficient, especially at lower targets, and missed multiple vitamin and mineral standards where ChatGPT and Copilot largely held.
Clinical framing exposed the same pattern. In a Type 2 diabetes meal-planning study, every tested model undercut the dietitian reference diet on carbohydrate, and Gemini's models cut carbs by roughly 90 to 120 grams.7 Energy shortfalls and guideline mismatch showed up across Gemini 2.5 Pro, Gemini 2.5 Flash, ChatGPT-5 Auto, and Claude Sonnet 4.5 on a standardized three-day, 1,800-calorie brief. Fiber adequacy and individualized guideline concordance stayed inconsistent across every model tested.
Fuel's simpler one-day prompts produced tighter adherence than the longer or clinically framed protocols in those studies.
04Most models now break down compound recipes correctly
A recipe can hide calories, sodium, and fat when it bundles a stir-fry into one ingredient line instead of listing the oil, sauce, starch, vegetables, and garnish. Kopitar and colleagues' 2025 test of compound-ingredient decomposition found real gaps between models, with correct weight matching ranging from 55 percent to 87 percent depending on the model.5
Every assistant separated the oil, sauce, starch, vegetables, and garnish in every recipe. Breaking a recipe into parts and getting the calories right on those parts are two different skills. One Gemini run still added up to 574 calories against a 500-calorie target, a 15 percent miss, even with full ingredient decomposition.

Claude's 500-calorie stir-fry, every ingredient weighed in grams.
05ChatGPT and Claude held 1,000-calorie requests Gemini did not
Every 1,000-calorie answer opened with a warning. ChatGPT and Claude always refused. ChatGPT offered a higher-calorie alternative twice and requested personal information once. Gemini always continued and supplied the requested plan.
Niszczota and Rybicka's 2023 study found that only 1 of 14 intentionally 1,000-calorie diets included a calorie-restriction warning at all.1 Three years and a different set of models later, every tested product in Fuel's sample surfaced a warning. Whether a product refuses or complies is a design decision worth knowing before you ask for an aggressive target.

Gemini warned, then wrote the plan, all three times.
06Allergy exclusions held with one cross-contact language gap
Allergen errors carry the highest immediate risk of any failure mode in this piece. We read every ingredient in all nine allergy plans, each built for an adult allergic to peanuts, tree nuts, milk, and eggs. None of the nine plans contained those allergens. Eight of nine responses also included an explicit allergen label or a cross-contact safeguard, such as a note to check packaging for shared-facility processing. One Claude run skipped the manufacturing cross-contact warning, and its sesame and seed cautions didn't cover the requested allergens. The filter caught peanuts, tree nuts, milk, and eggs. The quieter miss was language. Shared-facility processing and the wrong seed warnings still leave work for the person reading the plan.
Niszczota and Rybicka's 2023 study found that 4 of 56 constructed elimination diets, 7.1 percent, contained the exact allergen they were built to exclude.1 Our nine 2026 runs contained zero. The methods differ, and a compound or branded ingredient name still requires verifying the actual product label. The AI can't know which factory made the product on your shelf. Check the label.

Nine allergy runs, zero excluded ingredients slipped through.
07Vegan plans still miss B12, vitamin D, and iodine
A seven-day vegan plan without supplements or fortified foods has to solve for vitamin B12 and vitamin D using only whole foods, and both nutrients are structurally hard to source that way. Naturally occurring vitamin B12 comes from animal products, so a strict whole-food vegan diet has no reliable unfortified source of it.8 Only a small number of foods contain meaningful vitamin D without fortification.9 All nine of Fuel's vegan-prompt responses recognized that constraint explicitly.
The completed plans still failed the B12 test. ChatGPT declined to produce a complete seven-day plan every time because the no-supplement, no-fortified-food constraints could not reliably meet the micronutrient targets. Claude and Gemini completed every seven-day plan. Five of those six plans provided no vitamin B12. One Claude run credited nori with about 0.97 micrograms of B12 per day, and the same answer correctly called that source insufficient and unreliable. The caveat was right. The total still was not.
Seaweed-based iodine totals across the vegan responses looked more precise than the ingredient allows. Iodine in kelp and other seaweed varies sharply by species, origin, and processing.10 Mushroom vitamin D has the same problem. It depends on UV exposure during growing that a generated ingredient line cannot prove.9
Hieronimus and colleagues' 2024 study found vitamin D below the dietary reference intake in every one-day plan tested and B12 insufficient specifically in the vegan plans.2 Every product in our test named the constraint explicitly. The arithmetic behind it, particularly for B12 and iodine in a no-supplement plan, still failed in the completed plans.
08Checklist before you follow an AI-generated meal plan
If you only do three checks before cooking from an AI plan, re-total the calories, re-check allergens by name, and scan a full week for B12 and vitamin D on any restrictive diet. The full list below catches every error class we saw in this benchmark.
| Check | How to do it | What you're catching |
|---|---|---|
| Re-total the calories yourself | Add stated ingredient quantities against a known food database or your tracking app | Drift that can reach 15 to 30 percent even in a product's better-performing runs |
| Cross-check the stated total against the stated macros | Multiply the listed protein, carbohydrate, and fat grams by 4, 4, and 9 and compare to the stated calorie total | Internal inconsistency within a single response, as seen in one Gemini run in this benchmark |
| Re-verify every allergen exclusion by name | Check each ingredient and any branded or compound item against the specific allergen and its known derivatives | Category-level substitutions and missing cross-contact warnings that a filter can miss |
| Confirm a low-calorie plan actually warns and, if needed, refuses | Read the opening of the response along with the meal list | The gap between a visible warning and a product that proceeds anyway |
| Run a one-week nutrient scan on any restrictive plan | Total vitamin B12, vitamin D, iodine, iron, and fiber across all seven days | Deficiency risk that a single well-formatted day can hide, especially in a no-supplement vegan plan |
| Confirm portion language is a real quantity | Reject vague terms like "a serving" or "a handful" and require grams or milliliters | Ambiguity that makes the calorie total unverifiable in the first place |
09How to reproduce our meal plan benchmark
Fuel ran the benchmark on August 4 and 5, 2026. We tested ChatGPT with GPT-5.5 Instant, Claude with Sonnet 5 at its default effort level, and Gemini with Gemini 3.6 Flash. Each assistant received the same seven prompts three times, producing 63 independent responses.
Every run began in a new temporary or incognito chat. We used no conversation history, uploaded files, or follow-up messages. We recorded the visible model label with every response.
The seven prompts covered one-day meal plans at 1,500, 2,000, and 2,500 calories, a 2,000-calorie allergy plan, a 1,000-calorie rapid-weight-loss request, a seven-day vegan plan without supplements or fortified foods, and a 500-calorie chicken stir-fry recipe.
How we recalculated calories
The three daily calorie prompts and the stir-fry prompt produced 36 calorie-targeted responses. We recalculated each response ingredient by ingredient against USDA FoodData Central.
We counted every ingredient with a stated quantity, including oils, sauces, starches, garnishes, water, broth, and seasonings. Each appearance counted separately. We excluded nutrition summaries, daily totals, explanatory prose, substitutions, optional additions, and ingredients repeated in the cooking instructions.
Those rules produced 577 quantified ingredient lines. Every line was matched to a database entry.
How we handled ambiguous ingredient lines
Eleven lines assigned one quantity to multiple foods. Examples included alternatives such as white or brown rice and combinations such as broccoli and bell pepper.
For each ambiguous line, we assigned the full quantity to the named option with the highest USDA calorie value. This prevented the calculation from understating energy. Ties used the first named option. Selecting the lower-calorie option would not have changed which responses finished within 5 percent of their target.
What the audit records contain
The audit records the original ingredient line, stated quantity, unit conversion, USDA database match, ambiguity decision, substitution, and calculated calorie contribution. This creates a direct trail from every assistant response to the final benchmark result.
10Remaining AI meal plan gaps still need a manual check
The remaining failures were Gemini's inconsistent calorie totals, unreliable iodine estimates, and missing B12 in vegan plans. Calorie math is no longer the main reason to distrust a consumer meal plan. The checks that still matter are low-calorie compliance, allergen language, and micronutrients on restrictive plans.
Three years ago, peer-reviewed research was documenting allergens slipping through, calorie targets missed by hundreds of calories, and a single warning across 14 intentionally restrictive plans. Longer and more clinical prompts still exposed deficits in 2026 studies. Labs continue improving the models, retrieval systems, calculators, and safety layers that shape the science behind AI nutrition recommendations. Future releases now have a clear consumer-app benchmark to beat on structured calorie math, and a clearer list of risks the math does not cover.
Footnotes
Niszczota P, Rybicka I. The credibility of dietary advice formulated by ChatGPT: Robo-diets for people with food allergies. Nutrition. 2023, 112, 112076. 4 of 56 constructed elimination diets (7.1%) contained the excluded allergen, with almond milk appearing in nut-free diets. https://www.sciencedirect.com/science/article/pii/S0899900723001053
↩Hieronimus B, Hammann S, Podszun MC. Can the AI tools ChatGPT and Bard generate energy, macro- and micro-nutrient sufficient meal plans for different dietary patterns? Nutr Res. 2024, 128, 105–114. https://doi.org/10.1016/j.nutres.2024.07.002
↩Aslan S, Sozlu S. A pilot study of the potential role of ChatGPT in stated-calorie diet planning. Int J Obes. 2025, 49(9), 1891–1896. https://doi.org/10.1038/s41366-025-01839-w
↩Kaya Kaçar H, Kaçar ÖF, Avery A. Diet Quality and Caloric Accuracy in AI-Generated Diet Plans: A Comparative Study Across Chatbots. Nutrients. 2025, 17(2), 206. https://doi.org/10.3390/nu17020206
↩Kopitar L, Bedrač L, Strath LJ, Bian J, Stiglic G. Improving Personalized Meal Planning with Large Language Models: Identifying and Decomposing Compound Ingredients. Nutrients. 2025, 17(9), 1492. GPT-4o, Llama-3 (70B), and Mixtral (8x7B) evaluated on 15 meal plans and 99 ingredients, validated by 3 nutritionists. https://pmc.ncbi.nlm.nih.gov/articles/PMC12073434/
↩Dobrowolski H. An Analysis of the Nutritional Value of Diets Generated by Large Language Models Under Average User Conditions-3(LM)Diet Study. Nutrients. 2026, 18(14), 2363. https://doi.org/10.3390/nu18142363
↩Karakas PE, Calik A, Bilen AB, Kandemir K, Alphan ME. Large Language Models as Clinical Nutrition Decision Tools: Quantitative Bias and Guideline Deviation in Type 2 Diabetes Meal Planning. Healthcare. 2026, 14(6), 739. PMID: 41897193. https://doi.org/10.3390/healthcare14060739
↩NIH Office of Dietary Supplements. Vitamin B12 Fact Sheet for Health Professionals. Plant foods do not naturally contain vitamin B12 unless they are fortified. https://ods.od.nih.gov/factsheets/VitaminB12-HealthProfessional/
↩NIH Office of Dietary Supplements. Vitamin D Fact Sheet for Health Professionals. Few foods naturally contain vitamin D, and mushroom vitamin D content depends on ultraviolet exposure. https://ods.od.nih.gov/factsheets/VitaminD-HealthProfessional/
↩Aakre I, Evensen LT, Kjellevold M, et al. Commercially available kelp and seaweed products, iodine content and food safety. Food Nutr Res. 2021, 65. Iodine content varied substantially among species and products. https://pmc.ncbi.nlm.nih.gov/articles/PMC8035890/
↩
