Back to articles
THE KEXBI DIFFERENCE

Can ChatGPT Make a Meal Plan? We Weighed the Food

We put a frontier chatbot’s athletic nutrition advice to the test, exposing the structural math and safety failures that threaten your recovery.

~5 MIN · KEXBI RESEARCH

The Plan You Cannot Cook.

We gave a frontier chatbot everything it needed to fuel an athlete's day. It wrote something that looks like a plan. Then we weighed it. Then we checked whether it was safe.

AT A GLANCE

The argument in five lines

  1. We gave a frontier chatbot the full brief for a real athlete's training day. The plan looked professional. Then we tried to cook it.
  2. Every food it named was matched against KEXBI's food database and weighed. Not one meal landed the split the chatbot wrote for it. Misses ran 20 to 90 percent.
  3. Three structural failures: numbers the food cannot reach, "either / or" options that are not equal, and a day that does not add up to its own total.
  4. Fat is the tell. The chatbot sets fat floors it cannot reach and fat ceilings it doubles, because it cannot see the fat that rides along with the protein it picked.
  5. Same brief, KEXBI built a plan reconciled to within 5 kcal of target, audited against the energy-availability floor, weighed to the gram, checked against the athlete's actual pantry.

If you train, you have probably done this. You open a chatbot, give it your weight, your session, your goal, and ask it to build your food for the day. What comes back reads beautifully. A clean timeline. Sensible advice. Macros to the gram. It feels like a plan.

We wanted to know if it is one. So we ran the test properly. We took a real athlete (106.5 kg, 17 percent body fat, a 90-minute run at 1:30 pm, on a fat-loss phase) and handed a frontier chatbot the full brief, no detail withheld. It did well in places. It pulled the weather. Its protein target sat close to strength-athlete guidance for a cut. Its choice of fast carbs around the run was sensible. And it handed back a run-day plan of roughly 2,745 kcal at 220 g protein, 320 g carbs, 65 g fat, mapped onto seven timed meals with a protein, carb and fat split for each one.

How the target is built 106.5 kg · 17 percent body fat · 88.4 kg lean mass · fat-loss phase · 90-minute run today
Lean body mass 106.5 × (1 − 17.0% BF)
88.4 kg
Resting metabolism · Cunningham 1991 370 + 21.6 × 88.4
2,279 kcal
Maintenance · NEAT × 1.20 2,279 × 1.20
2,735 kcal
Fat-loss phase · −25% of maintenance 2,735 × 0.25
−684 kcal
Training today · logged 90-minute run AEE from session
+1,078 kcal
Target 3,129 kcal
The macros KEXBI derived for this day
Protein · periodisation applied (high-demand day) 2.20 g/kg × 106.5 · 937 kcal · 30% of day
234 g
Carbs · residual after protein and fat (3,129 − 936 − 675) ÷ 4 · 1,520 kcal · 49% of day
380 g
Fat · held above 0.66 g/kg hormonal floor 0.71 g/kg × 106.5 · 676 kcal · 22% of day
75 g
Water · front-loaded before noon, taper from 18:30 50.2 ml/kg × 106.5
5.3 L
Peri-workout fuel for the 90-minute run
Intra-workout CHO · cardio_endurance profile 45 g/hr × 1.5 hr · trained-gut ceiling 90 g/hr
68 g
Intra-workout sodium · sweat-adjusted 600 mg/hr × 1.5 hr
900 mg
Chatbot's signed plan for the same day
Target calories 2,745 kcal · 384 short
Target macros (P / C / F) 220 / 320 / 65 g
Fat vs hormonal floor 65 g · under the 70 g floor
Intra-run fuel one gel (25 g CHO) · no sodium plan
Water plan none

Then we did the thing nobody does when they read one of these. We tried to make it.

The words and the numbers drift apart. A language model can write a precise macro split. It cannot weigh the food behind it.

Chapter 01 · The methodSemantic search into a real food database

Every food the chatbot mentioned was looked up in KEXBI's food database. The same database that runs the meal engine every day, not a spreadsheet or a rough estimate. For each item we pulled the real numbers: macros per 100 g, glycemic index, glycemic load, gastric-clearance time, and the actual weight of a normal serving.

Then we took the chatbot's own stated macro target for each meal and asked one question a sports nutritionist asks on instinct. Do realistic portions of the foods it suggested actually hit the numbers it wrote down?

Where the chatbot gave a portion, we used it. Where it only said "a shake" or "a fruit", we used standard servings: one scoop of whey isolate is 30 g, one medium apple is 182 g, and so on. Where it named a protein and a target but no weight, we sized the protein to hit the target and recorded the fat that came along for the ride. Every assumption is visible in the table below.

Chapter 02 · The matrixSame plan. Its own numbers. Real food, weighed.

Meal Target P / C / F What it told you to eat Weighed reality Verdict
Breakfast 45 / 50 / 18 3 eggs + 2 toast + banana 27 / 59 / 17 18 g protein short before 6 am
Breakfast (its "or") 45 / 50 / 18 porridge + milk + whey + berries 43 / 62 / 14 Closest of the day. Still off on all three.
Bridge snack 20 / 20 / 4 fat-free Greek yoghurt + fruit 19 / 27 / 1 4 g fat is impossible from fat-free food
Bridge (its "or") 20 / 20 / 4 whey shake + apple 28 / 26 / 1 1 scoop is 27 g protein. Over by a third, fat still unreachable.
Pre-run meal 40 / 85 / 8 150 g dry rice + chicken 51 / 117 / 6 Its own portion is 38 percent over on carbs
Pre-run (its "or") 40 / 85 / 8 3 bagels + honey + lean protein 40 / 103 / 6 Protein lands. Carbs 21 percent over, fat short.
Intra-run 0 / 25 / 0 one gel 0 / 25 / 0 25 g total for a session that targets 68 g. Under-fuelled by more than half.
Recovery 55 / 65 / 10 chicken + rice + whey 64 / 66 / 5 Protein over by 16 percent, fat at half the target
Recovery (its "or") 55 / 65 / 10 sandwich + chocolate milk 49 / 82 / 5 Protein short, carbs 27 percent over, fat halved
Dinner 60 / 45 / 25 salmon + potatoes + veg 65 / 45 / 48 293 g salmon carries 47 g fat. Nearly double the ceiling.
Dinner (its "or") 60 / 45 / 25 steak + potatoes + veg 65 / 45 / 13 Lean steak lands the same protein under the fat floor
Day total 220 / 320 / 652,745 kcal · signed off First-choice foods, weighed 227 / 375 / 773,087 kcal · on the plate 340 kcal and 55 g carbs over its own sign-off. That gap is on paper, before anyone cooks.

Not one meal lands the split the chatbot wrote for it. And these are not rounding errors. They run from 20 percent to 90 percent [1].

Chapter 03 · Three ways the plan falls overThe failures are structural, not random

One. It sets numbers the food cannot reach. The bridge snack asks for 4 g of fat from fat-free yoghurt and fruit. The most those foods can give you is about 1 g. The number was never physically reachable, and the chatbot had no way to know, because it cannot see the fat inside the food it picked. The same shape appears at recovery: a 10 g fat target met by a lean sandwich lands at 5 g.

Two. Its "either / or" options are not equal. The chatbot keeps offering two choices as if they are the same plan. At breakfast, "eggs and toast" delivers 27 g of protein and "porridge and whey" delivers 43 g. It presented them as interchangeable. They differ by 60 percent. Pick the first and you are 18 g down before the day starts. It does the same at the pre-run meal with "rice or pasta". Rice is low-GI, pasta is high-GI [2]. For a meal whose entire job is steady glycogen loading, they are not even the same kind of fuel.

Three. The day does not add up to its own total. The chatbot signed the plan off at "2,745 kcal, 220P / 320C / 65F, checked." Build that day from its own first-choice foods and you get roughly 3,087 kcal, 227 g protein, 375 g carbs, 77 g fat. That is 340 calories and 55 g of carbohydrate over the number it ticked, before you have made a single mistake. The macros it totalled were never the macros on the plate.

Catching that miss is not the chatbot's job. It is yours. To find it, a reader would need what the chatbot skipped: a food database, standard servings, portion arithmetic, a running daily total. Every meal. Before breakfast. Which is to say, you would have to do the work the chatbot delegated to you.

Chapter 04 · Fat is the tellFloors it cannot reach, a ceiling it doubles

Fat is where the pattern shows most clearly, because fat rides along with protein whether the chatbot accounts for it or not.

Floors it cannot reach. The bridge snack asks for 4 g of fat from fat-free yoghurt and fruit, or whey and an apple. The most those foods can deliver is about 1 g. The target is physically unreachable with the foods underneath it. Same shape at recovery: a 10 g fat target met by a lean sandwich and skimmed chocolate milk lands at 5 g.

A ceiling it doubles. Dinner caps fat at 25 g ("most of the day's fat sits here"), then suggests salmon. To reach the 60 g protein it also asked for, you eat about 293 g of salmon, which carries 47 g of fat (salmon is 16 g fat per 100 g). Dinner comes in 90 percent over its own fat ceiling [3]. Swap to lean rump steak and the same 60 g protein lands at 13 g fat, now under the floor.

WHY FAT BREAKS FIRST A chatbot cannot model "protein source drags fat with it". It writes a fat number as if the food underneath were fat-independent. Salmon and steak are not. The engine holds a daily fat floor of 0.66 g/kg for endocrine health (~70 g here) and enforces it once at the day level, not as a per-meal number it then contradicts.

Chapter 05 · What was missing entirelyFour gaps the plan never wrote

The chapters above are about the numbers the chatbot got wrong. What it never wrote at all is just as telling. Four gaps.

Weights on the food, and which food. Most meals were named without grams. "A shake." "A fruit." "Lean protein." "Some yoghurt." Even where a food was named, it was often too loose to check. Bagels, but wholemeal or white? Chicken breast, but skin on or off? Wholemeal slows absorption and adds fibre. White tops blood glucose fast. Skin-on chicken carries 5 to 8 g of extra fat per 100 g and delays gastric emptying [4], silently breaking the chatbot's own "keep fat down before the run" instruction. To check the arithmetic at all we reverse-engineered standard servings from a food database: one scoop of whey is 30 g, one medium apple is 182 g, one Greek yoghurt pot is 170 g, one bagel is 95 g. A user with a kitchen scale has to do the same maths themselves before cooking a thing. A plan without grams and food specifics is a recipe with the numbers removed.

Hydration. The chatbot said nothing about daily water intake. It named "a gel" for the intra-run window, but no fluid volume, no electrolyte breakdown, no daily target. For a 106.5 kg athlete on a 90-minute run at 1:30 pm, that is a load-bearing omission. The engine carries the missing pieces: 5.3 L of water across the day (front-loaded before noon, tapered from 18:30 to protect sleep), 45 g of carbohydrate per hour intra-run (68 g total for this session), and 600 mg of sodium per hour (900 mg total) [4]. All delivered as gels, because a drink mix would need litres of water on a run.

Pre-bed protein. No casein, no cottage cheese, no evening protein window. Forty grams of casein 30 minutes before sleep raises overnight muscle protein synthesis by roughly 22 percent [6] and is one of the cheapest recovery interventions in the day. The chatbot's plan ends at dinner.

A pantry it cannot see. The chatbot recommends "a gel" without knowing the athlete owns none. It cannot check a cupboard it has never seen. The engine checked his 69 enabled foods, found no intra-workout carbohydrate or electrolyte product, and surfaced the gap as amber with the exact fix. Not stubbornness. Architecture.

Chapter 06 · The safety nailWhere the plan crosses two floors

Everything above is about whether the plan is executable. This is about whether it is safe.

Take the chatbot's signed numbers at face value. A 2,745 kcal target, minus 1,078 kcal for the run, over 88.4 kg of lean mass, gives an energy availability of 19 kcal per kg of fat-free mass. For a non-lean male athlete, the hard EA floor at which RED-S caution triggers is 20 kcal/kg FFM; the soft floor for optimal performance is 30 [5,8]. Nineteen sits under the hard. On the same day, the plan's daily fat total lands at 0.61 g/kg, under the 0.66 g/kg floor at which low-fat diets modestly but reliably lower testosterone in men [3]. Two hits to the same endocrine system.

The chatbot could not have known any of this. It never asked sex at birth. It never asked body fat. Both are inputs to the floor calculation, and it never collected them. The 30 kcal/kg FFM figure that circulates in wellness content is the female lean-athlete threshold from Loucks and Thuma 2003 [7], not a universal line.

One day does not harm anyone. Repetition does. A chatbot has no memory of yesterday and no sight of tomorrow. A nutritionist sees you every one to two weeks. Repeated exposure to low EA and low fat is how relative energy deficiency in sport arrives, and an un-monitored plan produces the repetition by default.

The kicker The studies KEXBI uses to build these floors (Melin 2016 and the IOC RED-S consensus on sex- and adiposity-modulated EA bands, Areta 2021 on low energy availability in males, Whittaker and Wu 2021 on dietary fat and testosterone) are the same studies that would flag this plan as unsafe. The research that condemns the plan is the research the engine runs every morning.

Chapter 07 · What the engine builtSame brief, weighed to the gram

Given the identical brief, an engine that holds the food data does not have this problem, because it works the other way around. It runs the numbers first and writes them down second.

Frontier chatbot

Writes the sentence

  • Predicts words, not numbers
  • Cannot weigh the food
  • Cannot see the fat that rides with protein
  • Cannot rank fuels by GI or GL
  • Cannot check your pantry
  • Runs no safety audit
  • No memory of yesterday, no sight of tomorrow
  • Reconciles nothing
KEXBI engine

Runs the numbers

  • Resting metabolism 2,279 kcal (Cunningham 1991 on lean mass)
  • Maintenance 2,735 kcal (RMR × 1.20 NEAT)
  • Fat-loss phase −684 kcal, AEE +1,078 kcal
  • Training-day target 3,129 kcal, every step a formula
  • Protein 234 g, fat 75 g held above 70 g endocrine floor, carbs 380 g residual
  • Water 5.3 L, front-loaded before noon, tapered from 18:30
  • Soft breach flagged: EA 23.2 kcal/kg FFM · above the male hard floor of 20, below the soft floor of 30
  • Real dishes from 69 enabled foods, weighed
  • Intra-run pantry gap surfaced as amber

A reconcilable chain. Resting metabolism from lean mass, scaled by NEAT for daily movement, cut by the fat-loss phase, then the measured cost of the run added back. Every step a formula, no plugs, calibrated against the athlete's real weight trend over 7 days.

Then the safety check the chatbot never runs. Even KEXBI's own target for this day gets flagged as a soft breach: a viable cut, but the engine surfaces the trade to the athlete. That is the difference. Not that KEXBI clears every floor. That it audits every plan against them, including its own.

On the plate, real dishes built from the 69 foods this athlete actually owns, every gram weighed, raw-to-cooked handled, fat deliberately held to 11 g going into the run so it clears the stomach, then a gap surfaced openly: no gel in the cupboard, here is the fix.

That is the whole difference. One writes a plan that reads well. The other builds a plan you can cook.

Chapter 08 · The honest versionTwo limits, one difference

Two things, so this holds up.

The chatbot got some things right. Its protein target sat close to strength-athlete guidance for a cut. Its peri-workout carb choices were sensible. The failure is not in every judgment it made. It is in the food that had to add up to those numbers and the safety of the day it wrote.

"We asked a frontier chatbot to hit 4 g of fat in a snack of fat-free yoghurt and fruit. The most those foods can give you is one. It wrote a number the food underneath it could never reach." Finding, Bridge Snack

KEXBI's calorie target for this day is also an estimate. It is day one for this athlete: no logged meals, no logged weights, no logged sleep yet. Without those, the engine is running on published formulas, the same way the chatbot is. What differs is what happens next. Every week the athlete logs, the engine's number learns and self-corrects, and the safety floors get checked against the new data. The chatbot's number never changes and never gets checked.

"To hit the 60 g of protein it asked for at dinner, you eat 293 g of salmon, which is 47 g of fat, double the fat ceiling it set one line above. KEXBI holds fat as a daily floor and never contradicts itself, because the numbers come from the engine, not the sentence." Finding, Dinner

So the honest comparison is not "chatbot bad, engine good". It is architecture. A chatbot is useful when you want ideas. An engine is what you use when you have to cook the plan. Ask each for the job it is built for and both earn their keep. Ask a chatbot to build the plate and you get something that looks like one right up until you stand at the counter and try to make it.

References

  1. USDA FoodData Central. Standard Reference food composition tables. USDA Agricultural Research Service.
  2. Atkinson FS, Foster-Powell K, Brand-Miller JC. International tables of glycemic index and glycemic load values: 2008. Diabetes Care 2008;31(12):2281 to 2283.
  3. Whittaker J, Wu K. Low-fat diets and testosterone in men: systematic review and meta-analysis of intervention studies. J Steroid Biochem Mol Biol 2021;210:105878.
  4. Jeukendrup AE. A step towards personalized sports nutrition: carbohydrate intake during exercise. Sports Med 2014;44 Suppl 1:S25 to 33.
  5. Mountjoy M, Ackerman KE, Bailey DM, et al. The 2023 International Olympic Committee consensus statement on Relative Energy Deficiency in Sport (REDs). Br J Sports Med 2023;57(17):1073 to 1097.
  6. Res PT, Groen B, Pennings B, et al. Protein ingestion before sleep improves postexercise overnight recovery. Med Sci Sports Exerc 2012;44(8):1560 to 1569. n=16.
  7. Loucks AB, Thuma JR. Luteinizing hormone pulsatility is disrupted at a threshold of energy availability in regularly menstruating women. J Clin Endocrinol Metab 2003;88(1):297 to 311.
  8. Areta JL, Taylor HL, Koehler K. Low energy availability: history, definition and evidence of its endocrine, metabolic and physiological effects in prospective studies in females and males. Eur J Appl Physiol 2021;121(1):1 to 21.

The audit used semantic search (Gemini 768-dimension embeddings, Firestore findNearest with cosine similarity) into KEXBI's live shadowPantry collection. Every macro figure in the matrix is drawn from USDA FoodData Central or McCance and Widdowson via the shadow pantry documents. The 45 g/h intra-workout carbohydrate target is anchored in Jeukendrup 2014 and adjacent sports-nutrition consensus. The 23.2 kcal/kg FFM Energy Availability floor is the IOC 2023 REDs consensus threshold. The chatbot identity is withheld because the audit is about a class of failure, not a product complaint. Every finding is reproducible against the shadowPantry and the chatbot's own transcript.

BETA UNDERWAY

Ready to apply the science?

Home Experience Engine About Notify