Can ChatGPT Make a Meal Plan? We Weighed the Food
We put a frontier chatbot’s athletic nutrition advice to the test, exposing the structural math and safety failures that threaten your recovery.
We put a frontier chatbot’s athletic nutrition advice to the test, exposing the structural math and safety failures that threaten your recovery.
We gave a frontier chatbot everything it needed to fuel an athlete's day. It wrote something that looks like a plan. Then we weighed it. Then we checked whether it was safe.
If you train, you have probably done this. You open a chatbot, give it your weight, your session, your goal, and ask it to build your food for the day. What comes back reads beautifully. A clean timeline. Sensible advice. Macros to the gram. It feels like a plan.
We wanted to know if it is one. So we ran the test properly. We took a real athlete (106.5 kg, 17 percent body fat, a 90-minute run at 1:30 pm, on a fat-loss phase) and handed a frontier chatbot the full brief, no detail withheld. It did well in places. It pulled the weather. Its protein target sat close to strength-athlete guidance for a cut. Its choice of fast carbs around the run was sensible. And it handed back a run-day plan of roughly 2,745 kcal at 220 g protein, 320 g carbs, 65 g fat, mapped onto seven timed meals with a protein, carb and fat split for each one.
Then we did the thing nobody does when they read one of these. We tried to make it.
Every food the chatbot mentioned was looked up in KEXBI's food database. The same database that runs the meal engine every day, not a spreadsheet or a rough estimate. For each item we pulled the real numbers: macros per 100 g, glycemic index, glycemic load, gastric-clearance time, and the actual weight of a normal serving.
Then we took the chatbot's own stated macro target for each meal and asked one question a sports nutritionist asks on instinct. Do realistic portions of the foods it suggested actually hit the numbers it wrote down?
Where the chatbot gave a portion, we used it. Where it only said "a shake" or "a fruit", we used standard servings: one scoop of whey isolate is 30 g, one medium apple is 182 g, and so on. Where it named a protein and a target but no weight, we sized the protein to hit the target and recorded the fat that came along for the ride. Every assumption is visible in the table below.
| Meal | Target P / C / F | What it told you to eat | Weighed reality | Verdict |
|---|---|---|---|---|
| Breakfast | 45 / 50 / 18 | 3 eggs + 2 toast + banana | 27 / 59 / 17 | 18 g protein short before 6 am |
| Breakfast (its "or") | 45 / 50 / 18 | porridge + milk + whey + berries | 43 / 62 / 14 | Closest of the day. Still off on all three. |
| Bridge snack | 20 / 20 / 4 | fat-free Greek yoghurt + fruit | 19 / 27 / 1 | 4 g fat is impossible from fat-free food |
| Bridge (its "or") | 20 / 20 / 4 | whey shake + apple | 28 / 26 / 1 | 1 scoop is 27 g protein. Over by a third, fat still unreachable. |
| Pre-run meal | 40 / 85 / 8 | 150 g dry rice + chicken | 51 / 117 / 6 | Its own portion is 38 percent over on carbs |
| Pre-run (its "or") | 40 / 85 / 8 | 3 bagels + honey + lean protein | 40 / 103 / 6 | Protein lands. Carbs 21 percent over, fat short. |
| Intra-run | 0 / 25 / 0 | one gel | 0 / 25 / 0 | 25 g total for a session that targets 68 g. Under-fuelled by more than half. |
| Recovery | 55 / 65 / 10 | chicken + rice + whey | 64 / 66 / 5 | Protein over by 16 percent, fat at half the target |
| Recovery (its "or") | 55 / 65 / 10 | sandwich + chocolate milk | 49 / 82 / 5 | Protein short, carbs 27 percent over, fat halved |
| Dinner | 60 / 45 / 25 | salmon + potatoes + veg | 65 / 45 / 48 | 293 g salmon carries 47 g fat. Nearly double the ceiling. |
| Dinner (its "or") | 60 / 45 / 25 | steak + potatoes + veg | 65 / 45 / 13 | Lean steak lands the same protein under the fat floor |
| Day total | 220 / 320 / 652,745 kcal · signed off | First-choice foods, weighed | 227 / 375 / 773,087 kcal · on the plate | 340 kcal and 55 g carbs over its own sign-off. That gap is on paper, before anyone cooks. |
Not one meal lands the split the chatbot wrote for it. And these are not rounding errors. They run from 20 percent to 90 percent [1].
One. It sets numbers the food cannot reach. The bridge snack asks for 4 g of fat from fat-free yoghurt and fruit. The most those foods can give you is about 1 g. The number was never physically reachable, and the chatbot had no way to know, because it cannot see the fat inside the food it picked. The same shape appears at recovery: a 10 g fat target met by a lean sandwich lands at 5 g.
Two. Its "either / or" options are not equal. The chatbot keeps offering two choices as if they are the same plan. At breakfast, "eggs and toast" delivers 27 g of protein and "porridge and whey" delivers 43 g. It presented them as interchangeable. They differ by 60 percent. Pick the first and you are 18 g down before the day starts. It does the same at the pre-run meal with "rice or pasta". Rice is low-GI, pasta is high-GI [2]. For a meal whose entire job is steady glycogen loading, they are not even the same kind of fuel.
Three. The day does not add up to its own total. The chatbot signed the plan off at "2,745 kcal, 220P / 320C / 65F, checked." Build that day from its own first-choice foods and you get roughly 3,087 kcal, 227 g protein, 375 g carbs, 77 g fat. That is 340 calories and 55 g of carbohydrate over the number it ticked, before you have made a single mistake. The macros it totalled were never the macros on the plate.
Catching that miss is not the chatbot's job. It is yours. To find it, a reader would need what the chatbot skipped: a food database, standard servings, portion arithmetic, a running daily total. Every meal. Before breakfast. Which is to say, you would have to do the work the chatbot delegated to you.
Fat is where the pattern shows most clearly, because fat rides along with protein whether the chatbot accounts for it or not.
Floors it cannot reach. The bridge snack asks for 4 g of fat from fat-free yoghurt and fruit, or whey and an apple. The most those foods can deliver is about 1 g. The target is physically unreachable with the foods underneath it. Same shape at recovery: a 10 g fat target met by a lean sandwich and skimmed chocolate milk lands at 5 g.
A ceiling it doubles. Dinner caps fat at 25 g ("most of the day's fat sits here"), then suggests salmon. To reach the 60 g protein it also asked for, you eat about 293 g of salmon, which carries 47 g of fat (salmon is 16 g fat per 100 g). Dinner comes in 90 percent over its own fat ceiling [3]. Swap to lean rump steak and the same 60 g protein lands at 13 g fat, now under the floor.
The chapters above are about the numbers the chatbot got wrong. What it never wrote at all is just as telling. Four gaps.
Weights on the food, and which food. Most meals were named without grams. "A shake." "A fruit." "Lean protein." "Some yoghurt." Even where a food was named, it was often too loose to check. Bagels, but wholemeal or white? Chicken breast, but skin on or off? Wholemeal slows absorption and adds fibre. White tops blood glucose fast. Skin-on chicken carries 5 to 8 g of extra fat per 100 g and delays gastric emptying [4], silently breaking the chatbot's own "keep fat down before the run" instruction. To check the arithmetic at all we reverse-engineered standard servings from a food database: one scoop of whey is 30 g, one medium apple is 182 g, one Greek yoghurt pot is 170 g, one bagel is 95 g. A user with a kitchen scale has to do the same maths themselves before cooking a thing. A plan without grams and food specifics is a recipe with the numbers removed.
Hydration. The chatbot said nothing about daily water intake. It named "a gel" for the intra-run window, but no fluid volume, no electrolyte breakdown, no daily target. For a 106.5 kg athlete on a 90-minute run at 1:30 pm, that is a load-bearing omission. The engine carries the missing pieces: 5.3 L of water across the day (front-loaded before noon, tapered from 18:30 to protect sleep), 45 g of carbohydrate per hour intra-run (68 g total for this session), and 600 mg of sodium per hour (900 mg total) [4]. All delivered as gels, because a drink mix would need litres of water on a run.
Pre-bed protein. No casein, no cottage cheese, no evening protein window. Forty grams of casein 30 minutes before sleep raises overnight muscle protein synthesis by roughly 22 percent [6] and is one of the cheapest recovery interventions in the day. The chatbot's plan ends at dinner.
A pantry it cannot see. The chatbot recommends "a gel" without knowing the athlete owns none. It cannot check a cupboard it has never seen. The engine checked his 69 enabled foods, found no intra-workout carbohydrate or electrolyte product, and surfaced the gap as amber with the exact fix. Not stubbornness. Architecture.
Everything above is about whether the plan is executable. This is about whether it is safe.
Take the chatbot's signed numbers at face value. A 2,745 kcal target, minus 1,078 kcal for the run, over 88.4 kg of lean mass, gives an energy availability of 19 kcal per kg of fat-free mass. For a non-lean male athlete, the hard EA floor at which RED-S caution triggers is 20 kcal/kg FFM; the soft floor for optimal performance is 30 [5,8]. Nineteen sits under the hard. On the same day, the plan's daily fat total lands at 0.61 g/kg, under the 0.66 g/kg floor at which low-fat diets modestly but reliably lower testosterone in men [3]. Two hits to the same endocrine system.
The chatbot could not have known any of this. It never asked sex at birth. It never asked body fat. Both are inputs to the floor calculation, and it never collected them. The 30 kcal/kg FFM figure that circulates in wellness content is the female lean-athlete threshold from Loucks and Thuma 2003 [7], not a universal line.
One day does not harm anyone. Repetition does. A chatbot has no memory of yesterday and no sight of tomorrow. A nutritionist sees you every one to two weeks. Repeated exposure to low EA and low fat is how relative energy deficiency in sport arrives, and an un-monitored plan produces the repetition by default.
Given the identical brief, an engine that holds the food data does not have this problem, because it works the other way around. It runs the numbers first and writes them down second.
A reconcilable chain. Resting metabolism from lean mass, scaled by NEAT for daily movement, cut by the fat-loss phase, then the measured cost of the run added back. Every step a formula, no plugs, calibrated against the athlete's real weight trend over 7 days.
Then the safety check the chatbot never runs. Even KEXBI's own target for this day gets flagged as a soft breach: a viable cut, but the engine surfaces the trade to the athlete. That is the difference. Not that KEXBI clears every floor. That it audits every plan against them, including its own.
On the plate, real dishes built from the 69 foods this athlete actually owns, every gram weighed, raw-to-cooked handled, fat deliberately held to 11 g going into the run so it clears the stomach, then a gap surfaced openly: no gel in the cupboard, here is the fix.
That is the whole difference. One writes a plan that reads well. The other builds a plan you can cook.
Two things, so this holds up.
The chatbot got some things right. Its protein target sat close to strength-athlete guidance for a cut. Its peri-workout carb choices were sensible. The failure is not in every judgment it made. It is in the food that had to add up to those numbers and the safety of the day it wrote.
KEXBI's calorie target for this day is also an estimate. It is day one for this athlete: no logged meals, no logged weights, no logged sleep yet. Without those, the engine is running on published formulas, the same way the chatbot is. What differs is what happens next. Every week the athlete logs, the engine's number learns and self-corrects, and the safety floors get checked against the new data. The chatbot's number never changes and never gets checked.
So the honest comparison is not "chatbot bad, engine good". It is architecture. A chatbot is useful when you want ideas. An engine is what you use when you have to cook the plan. Ask each for the job it is built for and both earn their keep. Ask a chatbot to build the plate and you get something that looks like one right up until you stand at the counter and try to make it.
The audit used semantic search (Gemini 768-dimension embeddings, Firestore findNearest with cosine similarity) into KEXBI's live shadowPantry collection. Every macro figure in the matrix is drawn from USDA FoodData Central or McCance and Widdowson via the shadow pantry documents. The 45 g/h intra-workout carbohydrate target is anchored in Jeukendrup 2014 and adjacent sports-nutrition consensus. The 23.2 kcal/kg FFM Energy Availability floor is the IOC 2023 REDs consensus threshold. The chatbot identity is withheld because the audit is about a class of failure, not a product complaint. Every finding is reproducible against the shadowPantry and the chatbot's own transcript.