← All answers

Taking action

How can I tell whether a small experiment is stable enough to evaluate?

A small experiment is stable enough for a modest review when you can name the version you tried, the action was carried out consistently enough to interpret, the planned evidence was recorded in a usable way, and no unlogged change or major context shift makes the attempts unlike one another. There is no universal number of attempts that proves stability. If the check fails, review implementation rather than outcomes, repair one issue, and begin a clearly labelled version. The STEADY check keeps that decision visible without turning personal observations into causal proof.

Match stability to the question

First write the smallest question the review must answer. “Did I carry out the planned version?” is an implementation question. “What helped or blocked it?” is a context question. “Did this action cause the outcome?” is a much stronger question that a few personal observations ordinarily cannot answer. Stability is therefore not one fixed standard; it is whether the experiment and its records are steady enough for the claim you intend to make.

CDC distinguishes questions about fidelity and implementation from questions about whether an intervention caused an outcome. The 2026 Test and Learn guidance likewise reserves stronger evaluation for a design that is sufficiently defined and operating consistently. For a small personal experiment, keep the conclusion proportionate: describe what happened in this version and what to repair, repeat, or reconsider next.

Stable enough for a practical review is not the same as strong enough for a causal conclusion.

Use the STEADY readiness check

STEADY is a Forever Free Compass framework, not an external research finding. Complete it at the planned review point before interpreting an encouraging or disappointing result.

  1. Same version: name the exact action, timing, support, and boundary that defined this version.
  2. Tried as specified: mark which planned opportunities happened as intended, which were missed, and why.
  3. Evidence usable: confirm the chosen note, count, observation, or feedback prompt was recorded consistently enough to answer the question.
  4. Alterations separated: move observations after a meaningful change into a new labelled version instead of pooling them with the old one.
  5. Disruptions named: record relevant context changes such as access, workload, timing, setting, or another person's availability.
  6. Yield matched to the question: state the narrow conclusion these observations can support and the stronger conclusion they cannot.

Check the version before counting the observations

Do not begin with an arbitrary attempt count. Two carefully recorded opportunities can reveal a delivery problem; many inconsistent attempts can still be hard to interpret. Ask whether the key action stayed recognisably the same, whether the opportunities were comparable enough for your question, and whether a planned change split the experiment into separate versions.

The 2026 Test and Learn readiness checklist looks for a stable design, consistent delivery, functioning outcome measures, and no major planned disruption before robust evaluation. Those criteria were written for public programmes, not personal choices. STEADY adapts only the underlying discipline of naming the version and checking consistency before interpreting it.

Make sure the evidence process worked

A stable action with unreliable notes is not ready for the question you planned. Check for missing entries, changed prompts, delayed recording, inconsistent definitions, and records that do not actually match the learning question. Explain discrepancies rather than quietly discarding them. If the method changed, preserve the earlier records and label when the new method began.

CDC's evidence guidance asks whether collection protocols were followed consistently, whether discrepancies were resolved or explained, whether collection was timely, and whether measures remained relevant. It also recommends documenting changes to a collection plan. The modest personal adaptation is a transparent record, not a claim of research-grade data quality.

If STEADY fails, review implementation first

A failed readiness check is useful information. Name the single largest interpretation problem: the action kept changing, opportunities were missed, the evidence method failed, or context shifted materially. Decide whether to repair that issue, start a new labelled version, or stop because the experiment is no longer feasible or appropriate. Keep any earlier observation attached to its original version.

The Magenta Book treats design, implementation, and outcomes as distinct parts of evaluation and recommends proportionate, fit-for-purpose work. CDC notes that observational approaches can answer noncausal questions about fidelity and implementation. When stability is weak, the honest output is therefore an implementation finding—not a verdict that the underlying direction works or fails.

What failedWhat you can sayNext move
The action changed between attemptsThe versions are not directly comparableSeparate the versions and test one defined version
Several planned opportunities were missedDelivery was inconsistent in this periodFind the delivery barrier before judging the outcome
The evidence record changed or has gapsThe planned measure did not function reliablyRepair or simplify the method and label the restart
A major context condition shiftedThe observations belong to different conditionsRecord the disruption and narrow or postpone the review

Keep the scope low-stakes and reversible

Use STEADY only to support a bounded next-step review. Do not use it to bypass consent, workplace or legal requirements, clinical or safety guidance, or another person's authority over their own choice. A consistent personal record does not show that an action will work for someone else, remain effective over time, or outweigh material risks.

If the consequences are significant, the evidence is sensitive, another person could be harmed or pressured, or a causal claim matters, use the required formal process or qualified guidance. The right conclusion may be that this informal experiment is not an appropriate way to answer the question.

Original contribution

The Forever Free Compass take

Forever Free Compass uses STEADY as an original review-readiness check: Same version; Tried as specified; Evidence usable; Alterations separated; Disruptions named; Yield matched to the question. STEADY makes the interpretation limit part of the check: a stable low-stakes personal experiment may support a practical decision about the next attempt, but it does not by itself establish causality, generalisability, effectiveness, or safety. It is not a validated evaluation method, statistical threshold, fidelity measure, research protocol, or substitute for consent, required processes, formal evaluation, or qualified guidance.

Sources

  1. Test and LearnHM Treasury and Evaluation Task Force. Stable design, consistent delivery, functioning measures, documented changes, readiness for robust evaluation, proportional methods, and cautious interpretation of early evidence.
  2. CDC Program Evaluation Framework, 2024Centers for Disease Control and Prevention. Distinguishing implementation and fidelity questions from causal outcome questions; evidence expectations, indicators, data quality, collection, and context.
  3. Step 4 – Gather Credible EvidenceCenters for Disease Control and Prevention. Question-aligned evidence; consistency, timeliness, relevance, completeness, discrepancies, data collection protocols, and documentation of method changes.
  4. Magenta Book: Central Government guidance on evaluationHM Treasury and Evaluation Task Force. Proportionate and fit-for-purpose evaluation; separate attention to intervention design, implementation, outcomes, monitoring, and context.

This is educational reflection, not medical or mental-health care. If distress is persistent, severe, or involves immediate safety, contact a qualified professional or local emergency service.

Run the STEADY check before interpreting one low-stakes experiment.

Start your free compass