← Back to Learn

Field note 11 · AI evidence

Before You Follow an AI Insight: Check the Evidence

Six questions, three worked examples and a downloadable card for checking the reasoning behind an app recommendation.

A fluent explanation still needs a chain of evidence

An AI insight can sound helpful while skipping the step that matters most: connecting a recommendation to the observations actually available. “Your sleep improved, so keep doing this” contains at least three separate claims. Something changed in the record; a particular habit explains it; continuing the habit will help. The first may be observable while the other two remain untested.

You do not have to evaluate a whole model to inspect one recommendation. Start with the evidence it names, the dates it covers and the action it proposes. The useful outcome is a smaller, more defensible statement—not a new numerical score for whether the AI is safe.

Nutrinaut’s own technical review helped shape the six questions below. We found stored fields for evidence, confidence, comparison days and recommendation text. That structure makes some checks possible. It does not make the recommendations correct, and our review is not a clinical validation of the product.

The six-question evidence card

  1. What observations were used? Look for the actual measurements or logged events, with units. “Your data suggests” is not enough to tell you whether the source was a watch reading, a manually entered tag or another calculated score.
  2. Which period do they describe? Separate the date an insight was generated from the dates of its inputs. A new message can summarise old records. Ask whether the comparison spans the same hours, days or weeks.
  3. What is missing or carried forward? A last-known value is not a new observation. Missing sleep, an incomplete food diary or an unlogged event changes what a comparison can establish.
  4. What is the comparison? Ask how many usable days are on each side, whether they belong to the same person and whether conditions were reasonably comparable. Repeated days are not independent participants.
  5. Is the explanation circular? If sleep is an input to a recovery score, finding that the score changes with sleep may partly describe the formula. It is not automatically an independent discovery about recovery.
  6. Does the action follow? Describing an association does not demonstrate its cause or show that changing a habit will improve an outcome. A stronger recommendation needs stronger evidence than a descriptive summary.

Keep the card beside the insight

Write down the source, event dates, missing information, comparison, possible circularity and proposed action. You can use the card privately; it does not require uploading personal records.

Two A4 pages for printing and private notes. The PDF uses the same method as the editable text version; it does not send your answers to Nutrinaut.

If a detail is unavailable, write “not shown” rather than guessing. That does not prove the model lacked the information internally. It means you cannot verify the explanation from what you have. The distinction keeps this checklist fair as well as cautious.

Three examples, three different conclusions

All examples are invented. Their numbers illustrate reasoning, not product performance or user outcomes.

1. Too little information to compare

Proposed insight: “On evenings with a late meal, your sleep is worse. Stop eating late.”

What is visible: two evenings have a late-meal tag. No evening is explicitly recorded as “no late meal,” and several nights have no sleep record.

Evidence check: two tagged events do not establish a comparison group. Untagged evenings may mean the event did not occur, or simply that nobody logged it. Missing nights could also change the summary. The instruction moves beyond what those records establish.

Defensible next step: define the event consistently, distinguish “no” from “unknown,” and inspect the available sleep records before interpreting a pattern. Do not change a health plan because this incomplete comparison sounds confident.

Smaller statement: “Two late-meal events are logged, but there is not enough verified comparison information here to assess a relationship.”

2. A useful descriptive observation

Visible records: three complete-looking nights from the same source show 7, 7.5 and 8 hours. Another three show 6, 6.5 and 7 hours. The time windows and source are documented.

First mean: 22.5 ÷ 3 = 7.5 hours
Second mean: 19.5 ÷ 3 = 6.5 hours

Defensible observation: “Recorded sleep averaged one hour longer in the first set of three nights.” The arithmetic is reproducible and the claim stays within the records.

What it does not establish: why the difference occurred, whether every minute was measured correctly, or whether a particular behaviour caused it. Three nights per group are not a general adequacy threshold. A source change or unusual week would still matter.

Useful action: check the dates and context, then retain the finding as a description. You have learned something about these records without pretending to have identified a treatment.

3. A recommendation that outruns its evidence

Proposed insight: “Your recovery score is higher on days after longer recorded sleep, so increase training intensity today.”

What is visible: the score’s documentation says sleep contributes to its calculation. No independent performance outcome or information establishing that today’s activity is appropriate is shown.

Evidence check: the association may partly reproduce the scoring rule. Even a genuine association would not, by itself, validate the proposed training change. The descriptive claim and the action require separate support.

Smaller statement: “This score partly reflects recorded sleep. That relationship does not independently determine how hard you should train today.”

Useful action: inspect the inputs and the product’s stated limits. Treat the score as context rather than an instruction that overrides symptoms, professional advice or an established plan.

Why “confidence: 0.8” is not automatically “80% correct”

A confidence field needs a definition. It might be a model-generated assessment, a rule-derived number or a score transformed for display. Without knowing what was measured and how it was evaluated, the decimal alone cannot tell you the probability that a recommendation is correct.

To support a probability interpretation, the developer would need to define a testable outcome and evaluate the score against appropriate outcomes. “Correct” is already a difficult word here: accurate arithmetic, faithful summarisation and a beneficial personal recommendation are different targets. Success at one does not establish the others.

Ask what the confidence refers to, which data were used for evaluation, and whether performance remains relevant to this kind of input. If that information is not available, keep the score as an unexplained system output. Do not convert it into a personal guarantee.

The NIST AI Risk Management Framework emphasises evaluation and the context in which a system is used. Citing a framework does not mean Nutrinaut has been independently assessed against it. Here it supports the need for explicit evaluation, not a product-quality claim.

What our own historical records let us check

On 29 September 2026, our reproducible inventory of 300 account nodes found 171 stored tag-insight records across three nodes. Their generation timestamps were in the first half of 2026, not the current September release. We inspected their structure and aggregate comparison fields; we did not independently verify the health effectiveness of 171 recommendations.

Of the 171 records, 120 had positive day counts for both comparison groups. In 96 of those 120, at least one group contained fewer than five days. This is a description of the stored evidence fields, not a failure rate. Five was a screening cutoff for our inspection, not a scientific boundary between adequate and inadequate evidence.

The selection was not a representative sample of active people, and the insight histories were heavily concentrated. We did not establish complete test-account classification, a current model version or independently reviewed outcomes. For those reasons, these historical records informed our questions—not a headline about current AI accuracy.

The data-reading guide offers a practical way to check units, coverage and conclusions. An honest audit should state what was checked and what remained outside its reach. Finding an evidence field is useful; independently testing what it claims is another task.

Turn the answer into a smaller decision

After filling the card, put the insight into one of three plain-language categories: “a description I can reproduce,” “a hypothesis needing better records,” or “an action not supported by the information shown.” These are reading aids, not validated safety classifications.

A descriptive observation can help you notice a logging problem or ask a focused question. A hypothesis can guide more consistent notes without being presented as a discovery. An unsupported action should not gain authority simply because it was phrased politely or accompanied by a confidence number.

For missing device records, use the source and date checklist. For habit comparisons, use the three-state tag journal. For meal estimates, verify the saved food and quantity. These are concrete checks you can perform before asking an AI system to interpret a larger pattern.

The best immediate result is often not a stronger recommendation. It is knowing exactly which part of an explanation you can verify, which part is uncertain and which action still needs an independent reason.

Sources and scope

  • Original Nutrinaut read-only technical inventory, checked 29 September 2026. A fixed-seed SHA-256 account-key ordering selected 300 nodes for inspection. Historical insight subset: first half of 2026; 171 records from three account nodes. Account nodes are storage units, not verified active people. This was not a transactionally frozen research cohort; it does not evaluate current AI effectiveness.
  • NIST AI Risk Management Framework.

This is an editorial evidence-reading guide, with AI-assisted analysis and drafting. It is not independent expert review, a calibrated risk instrument or proof of Nutrinaut’s health effectiveness.

Found a factual or arithmetic error? Use the contact page and include the article URL and the source that needs checking. See our corrections policy.