References · systems
Systems, planning and maintenance: the evidence
The research behind how MVIII structures your training week, the studies behind it, and how strong that evidence actually is. 11 peer-reviewed sources, last verified 25 August 2026.
Read this before trusting the numbers
This document grades higher than the one on how the coach talks, and it is worth knowing why. Its two strongest rows, SY-P-01 and SY-P-03, rest on 94 and 138 randomized tests respectively, and both measure whether people did the thing rather than how they felt about it. Where this document is strong, it is strong for the right reason. Where it is weak, it is weak in one specific place: almost nothing here was measured in a training app. Harkin's monitoring effect was found across 138 interventions of every kind, and its most useful moderator, that monitoring works better when the information is physically recorded, was established long before a phone did the recording. Sniehotta's coping planning was measured in cardiac rehabilitation. The inference to a fitness app is reasonable and it is still an inference. The single most misused finding in this whole subject area is SY-F-01. One study, 96 volunteers, and a range from 18 to 254 days to reach automaticity. It is universally quoted as "66 days", which is the median of a distribution so wide the median tells you almost nothing. It is stated here in full, with the range, because the rounded version is how a real finding becomes a slogan. That same study contains the finding MVIII's most-defended product decision has been missing: missing one opportunity to perform the behavior did not materially affect the habit formation process. See SY-P-05. It does not evidence that streaks are harmful. It does contradict the premise streaks rest on.
Grades: A multiple meta-analyses or systematic reviews in agreement. B one meta-analysis, or several consistent controlled trials. C limited or single trials, wide intervals, high heterogeneity, or cross-sectional, retrospective, biomechanical or survey designs only. — no systematic review or controlled trial was located. A real gap, stated rather than filled in.
The entries
SY-P-01The session is planned in advance, with a when and a where
PrescriptionTraining is scheduled to a specific day rather than intended in general. The app writes the plan before the week starts, so the decision of whether to train is not made fresh each morning. Band: min "a stated intention to train", standard "a named session on a named day", max "an if-then plan naming the time, the place, and the response to a specific obstacle".
Evidence gradeA
EffectAcross 94 independent tests, forming an implementation intention that specifies the when, where and how of goal striving in advance had a medium-to-large effect on goal attainment (d = .65). Implementation intentions promoted the initiation of striving, shielded ongoing pursuit from unwanted influences, aided disengagement from failing courses of action, and conserved capability for future striving. Component processes were supported: forming the plan both raised the accessibility of the specified opportunity and automated the response [1].
Population94 independent tests across goal domains.
Sources[1]
CaveatsMVIII programs the when and leaves the where to the member, so it sits between min and max on its own band. The obstacle half of the if-then structure is SY-P-02, which is a separate and weaker row. Nothing in [1] was measured in a training app.
SY-P-02The obstacle gets a plan too, which is what a make-up session is
PrescriptionWhen a session is missed, the app offers it back as a make-up rather than deleting it, and the member may dismiss it. Planning for the barrier is a distinct practice from planning the action, and it matters later in the process rather than earlier. Band: min "reschedule silently", standard "offer the missed session back, dismissible", max "prompt the member to name the obstacle and plan a response to it".
Evidence gradeB
EffectIn a longitudinal study of 352 cardiac patients followed to four months after discharge, action planning and coping planning were psychometrically distinct and operated differently across the change process. Action plans were more influential early; coping plans were more instrumental later. Participants with higher coping planning after discharge were more likely to report higher exercise levels at four months [2].
Population352 cardiac rehabilitation patients, longitudinal, self-reported exercise.
Sources[2]
CaveatsOne study, one clinical population, self-reported outcome, and no randomization to a planning condition. The timing finding is the useful part and is the reason MVIII's make-up offer is not front-loaded onto new members. Whether an app offering a missed session back functions as coping planning at all is untested. See SY-P-09.
SY-P-03Progress is recorded, not remembered
PrescriptionEvery set, session and outcome is written down by the app, and what the member has done is shown back to them. Recording is not bookkeeping, it is the intervention. Band: min "the member can review history on request", standard "progress is recorded automatically and surfaced without being asked for", max "progress is recorded, surfaced, and reported to another person".
Evidence gradeA
EffectAcross 138 studies (N = 19,951) randomly allocating participants to a progress-monitoring intervention or control, interventions raised the frequency of progress monitoring (d+ = 1.98, 95% CI [1.71, 2.24]) and promoted goal attainment (d+ = 0.40, 95% CI [0.32, 0.48]), with the change in monitoring frequency mediating the effect on attainment. Moderation tests found larger effects on attainment when outcomes were reported or made public, and when the information was physically recorded [3].
Population138 randomized studies across goal domains.
Sources[3]
CaveatsThe public-reporting moderator is the uncomfortable one. It is the strongest single argument in this document for a social feature, and MVIII refuses social features for reasons recorded in self-image-programming.md at SI-H-04. That is a trade the app makes knowingly, at a measurable cost, and this row is where the cost is written down rather than argued away. "Physically recorded" in a 2016 meta-analysis predates the phone doing it, and whether an automatic log carries the same benefit as a written one is not established.
SY-P-04The same behavior in the same context
PrescriptionSessions are scheduled to consistent days rather than scattered to fit the week, because automaticity is built by repetition in a stable context rather than by repetition alone.
Evidence gradeC
EffectNinety-six volunteers performed a chosen behavior daily in the same context for 12 weeks. Automaticity rose along an asymptotic curve modellable at the individual level, and performing the behavior more consistently was associated with better model fit [4]. Habits emerge from gradual learning of associations between responses and the features of performance contexts that have historically covaried with them; once formed, perception of the context triggers the response without a mediating goal [5].
Population96 volunteers, 82 with sufficient data, model fitting for 62 and a good fit for 39 [4]; theoretical model [5].
CaveatsThe habit-formation study [4] is a single study whose behaviors were mostly small daily acts, not training sessions, and the model fitted well for well under half the sample. The habit-goal model [5] is a model paper, not a trial. A training session is a poor candidate for pure habit formation: it is effortful, long, and not the kind of response a context cue triggers automatically. Read this row as support for a consistent schedule, not as a claim that training becomes automatic.
SY-P-05A missed session does not reset anything
PrescriptionNothing the member has earned is ever revoked, no counter falls to zero, and a missed session costs only that session. Consistency is measured in weeks trained, which only rises.
Evidence gradeC
EffectIn the only study located that measured the question directly, missing one opportunity to perform the behavior did not materially affect the habit formation process [4]. Consistent with this, habit associations accrue slowly and do not shift appreciably with infrequent counterhabitual responses [5].
Population96 volunteers over 84 days [4]; theoretical model [5].
CaveatsThis row does not evidence that streaks are harmful, and must not be read as though it does. No trial isolating streak mechanics in a training app was located, in either direction, and that gap is recorded at SI-P-12 in the companion document. What this row establishes is narrower and still useful: the premise that a broken chain undoes accumulated progress is contradicted by the one study that measured it. MVIII's design follows the evidence on that specific point and is a product decision on the rest.
SY-P-06Somebody six months in is not handled like somebody in week one
PrescriptionThe app distinguishes starting from continuing, and does not apply the same prompts, framing and structure to both.
Evidence gradeB
EffectA systematic review identified 117 behavior theories, of which 100 formulated hypotheses about maintenance. Five overarching interconnected themes emerged, and the review's central finding is that theoretical explanations for behavior change and for behavior change maintenance follow distinct patterns, differing in the nature and role of motives, self-regulation, psychological and physical resources, habits, and environmental and social influences from initiation to maintenance [6]. Behavior change interventions are effective at achieving temporary change; maintenance is rarely attained [6].
Population100 behavior theories, synthesized thematically with cross-validation.
Sources[6]
CaveatsA synthesis of theories, not of outcomes. It establishes that the field thinks initiation and maintenance are different, not that any particular way of handling them works. It is graded B on the strength and systematicity of the review rather than on trial evidence, and no row in this document should be read as a validated maintenance protocol.
SY-P-07What the app does is specified, not improvised
PrescriptionEvery behavioral technique MVIII uses is a named, deterministic thing in code rather than an emergent property of a model's phrasing, so it can be described, tested and cited.
Evidence gradeC
EffectA Delphi exercise with 14 experts rating 124 candidate techniques, an open-sort task with 18 further experts, and inter-rater agreement testing across six researchers coding 85 intervention descriptions, produced a taxonomy of 93 distinct behavior change techniques in 16 groups. Of the 26 techniques occurring at least five times, 23 had adjusted kappas of 0.60 or above [7].
PopulationExpert consensus exercise, not participants.
Sources[7]
CaveatsA reporting standard, not an effectiveness finding. It says nothing about whether any technique works. It is cited because it is the reason MVIII can state which techniques it uses at all, and because "we do behavior change" without naming the technique is the thing the taxonomy exists to stop.
SY-P-08Prompt and review cadence
PrescriptionMVIII prompts and surfaces progress on a cadence it chose. Nothing published establishes the right one.
Evidence grade—
EffectNo systematic review or controlled trial establishing an optimal frequency for an app to prompt a member or surface progress review, for training adherence, was located.
PopulationNone.
SourcesNone.
CaveatsSY-P-03 establishes that monitoring works. It says nothing about how often. The cadence is a design decision and is recorded as one.
SY-P-09Whether an offered make-up is coping planning
PrescriptionMVIII offers a missed session back and treats that as planning for the obstacle. Whether an app doing it on the member's behalf carries the benefit measured when the person does it themselves is untested.
Evidence grade—
EffectNo trial was located testing an automated make-up offer against either no offer or a member-authored coping plan. The rehabilitation study [2] measured coping planning that participants generated, which is not the same act.
PopulationNone.
SourcesNone.
CaveatsThis is the most consequential gap in the document, because the make-up mechanism is one of the few places MVIII's schedule differs from a plain calendar. The mechanism is reasonable and the specific claim is unevidenced. # PART B — Findings
SY-F-01How long a habit takes, stated properly
FindingNinety-six volunteers chose an eating, drinking or activity behavior to perform daily in the same context for 12 weeks, self-reporting automaticity each day. Nonlinear regression fitted an asymptotic curve to individual automaticity scores over 84 days. The time taken to reach 95% of individual asymptote ranged from 18 to 254 days [4].
Evidence gradeC
Population96 volunteers, 82 providing sufficient data, model fitting for 62, good fit for 39.
Sources[4]
CaveatsThis is one study, and the widely repeated "66 days" is a median of a range spanning more than eight months. The model did not fit well for over half the sample. The behaviors were small daily acts. Any product, book or coach quoting a single number here is quoting the middle of a distribution and discarding the finding, which is that the variation between people is enormous and it can take a very long time.
SY-F-02Nearly half of the people who intend to train do not
FindingA meta-analysis of 10 studies (N = 3,899) applying the action control framework at public-health activity guidelines estimated non-intenders who did not act (21%), non-intenders who acted anyway (2%), intenders who failed to follow through (36%), and successful intenders (42%). The overall intention-behavior gap was 46% [8].
Evidence gradeB
Population10 studies, N = 3,899, against public-health activity guidelines.
Sources[8]
CaveatsThe authors conclude this emphasizes the weakness of early intention models and would be a problem during intervention. It is the honest frame for every reminder MVIII sends: more than a third of people who fully intend to train will not, and that is the normal case rather than a failure of character. Only 10 studies met eligibility.
SY-F-03Self-regulation as a feedback loop
FindingThe discrepancy-reducing feedback loop, comparing a current state against a reference value and acting to reduce the difference, fits known phenomena across personality-social, clinical and health psychology and makes a distinct contribution in each [9].
Evidence gradeC
PopulationTheoretical synthesis.
Sources[9]
CaveatsA model paper from 1982, cited as the framework that SY-P-03's monitoring evidence sits inside, not as evidence of anything itself. It is the reason "monitor progress" is a mechanism rather than a habit of tidy people.
SY-F-04Advice about what to change wears off
FindingBrief advice on what to change and why engages deliberative processes, and the effects are typically short-lived because motivation and attention wane. Even where patients successfully initiate a change, gains are often transient, because few traditional behavior change strategies have built-in mechanisms for maintenance. Advice on how to change, engaging automatic processes, is proposed as the alternative with long-term potential [10].
Evidence gradeC
PopulationPractice-directed review, no primary sample.
Sources[10]
CaveatsA short review aimed at general practitioners, not a trial. It is cited for one sentence that describes MVIII's whole design premise better than anything else located: strategies without a maintenance mechanism produce transient gains. That is a hypothesis here, not a demonstration. # PART C — Harms
SY-H-01Goal setting overprescribed
FindingThe beneficial effects of goal setting have been overstated and the systematic harm it causes largely ignored. Identified side effects include a narrow focus that neglects non-goal areas, distorted risk preferences, a rise in unethical behavior, inhibited learning, corrosion of culture, and reduced intrinsic motivation. The authors argue goal setting should be treated as a prescription-strength medication requiring careful dosing and supervision rather than a benign over-the-counter treatment [11].
Evidence gradeC
PopulationReview and argument across the management goal-setting literature. Not exercisers.
Sources[11]
CaveatsA position piece in a management journal, contested within its own field, and graded accordingly. It is included because it is the counterweight to the goal-setting orthodoxy every fitness app is built on, and because "reduced intrinsic motivation" connects directly to the autonomous-motivation evidence in the companion document at SI-P-01. It is the stated basis for MVIII not attaching hard deadlines to body outcomes.
What this program will not tell you
Things commonly prescribed with confidence that the research does not currently support. MVIII programs none of them.
- That it takes 66 days to form a habit. One study, 96 volunteers, a range of 18 to 254 days, and a model that fitted well for 39 of them [4]. The number is a median presented as a constant.
- That breaking a chain of consecutive days undoes your progress. Directly contradicted by the only study that measured it: missing one opportunity did not materially affect habit formation [4], and habit associations do not shift appreciably with infrequent counterhabitual responses [5].
- That specific, challenging goals are reliably beneficial. The paradigm is heavily replicated and its side effects are systematically under-reported: narrowed focus, distorted risk preference, inhibited learning and reduced intrinsic motivation [11].
- That intending to train is most of the work. Thirty-six percent of people who intend to meet activity guidelines do not, for an overall gap of 46% [8].
- That a training session becomes automatic if you repeat it enough. The habit literature describes context-triggered responses learned gradually [4][5]. A ninety-minute effortful session is a poor fit for that mechanism, and no study located has shown training sessions becoming automatic in the sense the habit literature means.
- That an app recording your progress carries the same benefit as recording it yourself. Monitoring works, and the moderators favor information that is physically recorded and reported to others [3]. Neither condition describes an automatic phone log, and the difference is untested.
- That a make-up session offered by software is coping planning. The planning evidence measured plans participants generated themselves [2]. The substitution is untested and recorded at
SY-P-09. ---
References
- Gollwitzer PM, Sheeran P. Implementation Intentions and Goal Achievement: A Meta-analysis of Effects and Processes. Advances in Experimental Social Psychology. 2006. https://doi.org/10.1016/s0065-2601(06)38002-1
- Sniehotta FF, Schwarzer R, Scholz U, Schüz B. Action planning and coping planning for long-term lifestyle change: theory and assessment. Eur J Soc Psychol. 2005. https://doi.org/10.1002/ejsp.258
- Harkin B, Webb TL, Chang BPI, et al. Does monitoring goal progress promote goal attainment? A meta-analysis of the experimental evidence. Psychological Bulletin. 2016. https://doi.org/10.1037/bul0000025
- Lally P, van Jaarsveld CHM, Potts HWW, Wardle J. How are habits formed: Modelling habit formation in the real world. Eur J Soc Psychol. 2010. https://doi.org/10.1002/ejsp.674
- Wood W, Neal DT. A new look at habits and the habit-goal interface. Psychological Review. 2007. https://doi.org/10.1037/0033-295x.114.4.843
- Kwasnicka D, Dombrowski SU, White M, Sniehotta F. Theoretical explanations for maintenance of behaviour change: a systematic review of behaviour theories. Health Psychology Review. 2016. https://doi.org/10.1080/17437199.2016.1151372
- Michie S, Richardson M, Johnston M, et al. The Behavior Change Technique Taxonomy (v1) of 93 Hierarchically Clustered Techniques: Building an International Consensus for the Reporting of Behavior Change Interventions. Ann Behav Med. 2013. https://doi.org/10.1007/s12160-013-9486-6
- Rhodes RE, de Bruijn G. How big is the physical activity intention-behaviour gap? A meta-analysis using the action control framework. Br J Health Psychol. 2013. https://doi.org/10.1111/bjhp.12032
- Carver CS, Scheier MF. Control theory: A useful conceptual framework for personality-social, clinical, and health psychology. Psychological Bulletin. 1982. https://doi.org/10.1037/0033-2909.92.1.111
- Gardner B, Lally P, Wardle J. Making health habitual: the psychology of 'habit-formation' and general practice. Br J Gen Pract. 2012. https://doi.org/10.3399/bjgp12x659466
- Ordóñez LD, Schweitzer ME, Galinsky AD, Bazerman MH. Goals Gone Wild: The Systematic Side Effects of Overprescribing Goal Setting. Academy of Management Perspectives. 2009. https://doi.org/10.5465/amp.2009.37007999