S.M.M. Field Manual

Stop Manipulating Me · A Field Guide to Psychological Influence

FIELD MANUAL FM 1-6
SECTOR Learning · Conditioning
CLEARANCE Training Use
EDITION 01
Part I · Foundations Chapter
Operation · Silent Lever · Callsign ANVIL-06

Ch 06 — Learning & Conditioning

Behavioral Shaping · Behavior is shaped by its consequences, often below awareness
Conditioned

Mission Brief

Objective

Earlier chapters described systems that fire in a single moment; this one describes influence that works across time. Learning reshapes behavior through repeated pairings of cue, action, and consequence until it runs automatically, below deliberation. Conditioning is that process turned into a design discipline — cues placed where a designer wants them, rewards dosed on schedules chosen for compulsive power, behaviors built up increment by increment. Category 18 is this science operationalized, and it underlies "digital addiction," predatory monetization, and the compulsive core of coercive relationships.

Why It Matters

The analyst stops asking "why did I decide to do that?" — because often you didn't decide; you were trained — and starts asking "what cue triggered this, what schedule is rewarding it, and who arranged them?" This converts a felt compulsion from a personal failing into a diagnosable, interruptible design. Because conditioned behavior is built by repetition, it is unbuilt by structural changes to cues and rewards — the defenses are environmental, not willpower.

Sector
Behavioral Conditioning · Cat 18 (+ Cat 19)
Doctrine Ref
Mechanism §2.4 reward; §2.23 variable reinf.; §2.22 habit
Prerequisite
FM 1-5 §7 — loss aversion & emotional cycling
Confidence
C3–C5 (core findings)
Word Order
Analytic & defensive throughout
Operational Priority
Conditioning Layer
Reinforcement schedules install habits you never chose.
Governing DiagnosticChange the structure, not the willpower. Because the loop is built from cue, schedule, and friction, it is broken by changing cue, schedule, and friction — not by resisting the compulsion mid-routine. Willpower fights the routine at its strongest point; structure removes the cue before the routine ever fires.

1. Situation · The Terrain

Behavioral science distinguishes two learning engines, and manipulation uses both. Classical (Pavlovian) conditioning teaches that one stimulus predicts another — Pavlov’s dogs salivated at a bell that predicted food (Pavlov, 1927) [C5]; the Rescorla–Wagner refinement is that learning tracks surprise, not mere co-occurrence (Rescorla & Wagner, 1972) [C4]. So neutral cues acquire the charge of what they precede: a notification sound becomes arousing, a logo inherits positive affect, a partner’s tone triggers dread or longing before a word is processed. Operant (instrumental) conditioning teaches that one’s own behavior produces consequences — Thorndike’s Law of Effect (1911) [C5], formalized by Skinner into reinforcement and punishment (Skinner, 1938) [C5]. Control when and how often reward arrives and you control how durable and compulsive the behavior becomes.

Why unpredictable reward is most powerful

The single most important fact: how predictably a reward follows behavior matters more than its size (Ferster & Skinner, 1957) [C5]. Continuous reward builds fast but extinguishes fast; variable-ratio schedules — reward after an unpredictable number of actions — produce the highest, steadiest rates and the greatest resistance to extinction. This is the slot-machine schedule, deliberately engineered into feeds, matches, loot boxes, and trading apps (T18.1 ★, T18.5, T18.7). Skinner’s “superstition” pigeons developed rituals around accidentally-timed food (Skinner, 1948) [C4]; humans do the analogous refresh-tap rituals around any variable reward.

The brain rewards surprise; wanting ≠ liking

Midbrain dopamine encodes a reward prediction error — a fully predicted reward produces little response; an unexpected one produces a large one (Schultz, Dayan & Montague, 1997; Schultz, 2016) [C4]. A variable schedule maximizes prediction error, and thus the motivating signal (T18.21). Crucially, dopamine tracks wanting (incentive salience), dissociable from liking (Berridge & Robinson, 1998) [C4]: conditioning can sensitize wanting while liking flattens — the mechanistic signature of compulsion. “I don’t even enjoy it anymore” is a coherent description of a heavily conditioned habit. Repetition in a stable context then produces a habit — cue-triggered performance independent of current value (Wood & Rünger, 2016) [C4], the cue → routine → reward loop (Duhigg, 2012) [C3], with automaticity plateauing after a median ~66 days (Lally et al., 2010) [C3].

Table 6.1 — Reinforcement schedules and their manipulative use

ScheduleBehavioral effectWhere it appearsDefense
ContinuousFast to build, fast to extinguishOnboarding rewards, “beginner’s luck”Notice when payoff drops after you’re hooked
Fixed-ratioPost-reward pause, then push to next“Buy 9, get the 10th free”Evaluate the real value of the payoff
Fixed-intervalPause then acceleration near deadlineDaily login bonus, timed refillsIgnore the timer; act on actual need
Variable-ratioHighest rate; extinction-resistantFeeds, likes, matches, loot boxes, slotsScheduled not reactive use; add friction; name the slot-machine pattern
Variable-intervalSteady moderate checkingEmail, notificationsBatch into set windows

2. Enemy Forces · The Loop, Shaping & Predation

Design disciplines build products around the habit loop explicitly: Fogg’s model holds a behavior occurs when Motivation, Ability, and a Trigger converge (Fogg, 2009) [C3]; Eyal’s “Hooked” chains Trigger → Action → Variable Reward → Investment, where the final investment step raises switching costs and loads the next trigger (Eyal, 2014) [C3]. These are ethically neutral tools; Category 18 turns them against the user via cue engineering (notifications T18.2, “we miss you” T18.11), friction manipulation (infinite scroll T18.6, autoplay T18.12 remove stopping cues while exit friction is added), and reward & investment (variable-schedule likes T18.4, streaks T18.3 that convert quitting into a loss).

Shaping reinforces successive approximations (Skinner, 1938) [C5]: a manipulator rewards a small action, then a slightly larger one, so no single step is large enough to trigger refusal — the conditioning under foot-in-the-door (T6.1) and salami escalation (T6.9, T6.13). The gamification stack (T18.14–T18.18) is shaping formalized. The defensive key: compare every step to the original baseline, not the previous step. The purest predatory form is machine gambling (Schüll, 2012) [C3]: the near-miss recruits win-related reward circuitry despite being an objective loss (Clark et al., 2009) [C3], and loss-chasing escalates stakes. Loot boxes (T18.7) import variable-ratio, near-miss, and completion pressure into games sold to minors.

Table 6.2 — The manipulative habit loop, stage by stage

Loop stageManipulative deployment (Cat 18)DetectionDefense
Cue / triggerManufactured notifications, “we miss you” (T18.2, T18.11, T18.22)Cues serving the app’s schedule, not yoursDisable non-essential alerts; act on intent
Ability / frictionAutoplay, infinite scroll; friction on exit (T18.6, T18.12)No stopping point; hard to leaveRestore stopping cues; add friction to the routine
Variable rewardSlot-machine likes/matches/loot (T18.1, T18.4, T18.5, T18.7)Compulsive checking; wanting > likingScheduled use; recognize variable-ratio pull
InvestmentLock-in via data/followers/streaks (T18.3)Staying to avoid “losing” what you’ve put inWeigh sunk cost vs. real forward value

Table 6.3 — Conditioning techniques by domain

DomainTechnique (ID)DetectionDefense
Feeds / socialVariable rewards, likes (T18.1, T18.4)Compulsive checking; mood tied to metricsScheduled use; disable badges; decouple worth from metrics
RetentionStreaks, notifications (T18.3, T18.2, T18.11)Staying to preserve a streak; unprompted opensAllow breaks; disable alerts; value the activity not the streak
MonetizationLoot boxes, gambling mechanics (T18.5, T18.7)Paid random rewards; near-miss; loss-chasingHard limits; check odds; keep from minors
RelationshipsIntermittent reinforcement (T7.19, T22.11)Unpredictable warmth/coldness; eggshellsName the pattern; rebuild outside support; seek help

Relational conditioning (high-harm, recognition level)

Intermittent reinforcement in relationships (T18.9 → T7.19; T22.11) is the variable-ratio schedule enacted with affection, approval, and safety as the reward. Delivered unpredictably — interspersed with coldness or withdrawal — it produces the same extinction-resistant persistence a slot machine does, and the attachment strengthens precisely because the reward is unpredictable. This is the conditioning core of trauma bonding and coercive control: the intermittency forges the bond. Treated at recognition level only — the diagnostic is structural (unpredictable oscillation; wellbeing routed through regaining another’s warm state), and interrupting it typically requires stable outside support and often professional help.

Indicators · Warning OrderDetection · S-C-A-R-E-D
  • Cue-triggered opening. You opened an app without deciding to — a cue fired the routine automatically. [S-C-A-R-E-D: D · deviation from normal]
  • Compulsive checking for an uncertain payoff. More engagement when rewards are erratic than steady; a "just one more" pull; rituals around the action (T18.1). [R · reward lever scheduled for persistence]
  • Wanting outrunning liking. Strong urge with little enjoyment; anticipation far exceeding satisfaction; engagement rising while reward is flat or falling. [A · affect / conditioned pull]
  • No stopping point. Natural session-ending cues removed (infinite scroll, autoplay); leaving feels harder than arriving (T18.6, T18.12). [E · exit friction]
  • Escalation from baseline. Each step "just a little more"; your current normal is far from where you started (T6.9, T6.13). [C · consistency / commitment creep]
  • Unpredictable warmth-and-coldness. One person's oscillation, "walking on eggshells," working harder for intermittent affection; isolation from other support (T7.19). [E · isolation]
Rules of Engagement · Defense DoctrineP-A-U-S-E-D

Conditioning is not inherently manipulative — it is the mechanism of all learning, and used honestly builds skill, exercise, and saving habits. The marks of ethical conditioning: the reward is proportionate and real; cue and friction serve the user's own goals (easy to start and easy to stop); and the design survives disclosure. Manipulative conditioning engineers compulsion the user would not choose and depends on their not noticing a loop was installed. Because the loop is built from cue, schedule, and friction, the defensive posture is architectural.

ROE // Re-architect the Loop
  • Attack the cue. Habits are far easier to interrupt at the trigger than mid-routine. Disable non-essential notifications, remove habit-forming apps from the home screen, grayscale the display, keep devices out of the bedroom. Remove the cue and the routine has nothing to fire on.
  • Break the schedule. Convert reactive checking into scheduled checking (fixed windows for email, feeds, markets). Scheduling destroys the variable-ratio contingency that produces compulsion (P-A-U-S-E-D: Enforce process).
  • Restore stopping cues and re-friction the routine. Turn off autoplay and infinite scroll; set session timers; impose "one episode / one round" limits in advance. Raise friction on the routine and lower it on the exit.
  • Separate wanting from liking. Before acting on an urge, ask whether you'll actually enjoy it or merely quiet the pull. A large wanting-minus-liking gap means the behavior is conditioned, not chosen.
  • Anchor escalation to baseline & diversify reward. Measure "just a little more" against where it started, not the previous step (P-A-U-S-E-D: Pause the clock). Against intermittent-reinforcement control, rebuild predictable support from other relationships and seek professional help.
Field ChecklistLayer

Spot the Loop

  • Did I decide to do this, or did a cue trigger it automatically? What was the cue?
  • Is the reward variable/unpredictable? Am I checking more because the payoff is erratic (T18.1)? Any natural stopping point (T18.6, T18.12)?

Wanting vs. Liking

  • Do I actually enjoy this, or am I just quieting an urge? (Wanting > liking = conditioned.)
  • Is my engagement rising while my satisfaction is flat or falling?

Escalation, Lock-in & Predation

  • Is each step "just a little more"? Compared to where I started, how far has this gone (T6.9, T6.13)? Am I staying to avoid losing a streak (T18.3)?
  • Am I spending money/effort on chance-based rewards or seeing "so close" near-misses (T18.5, T18.7)? Is a spend mechanic aimed at a minor?

Relational & Structural

  • Is one person's warmth unpredictable — hot then cold — am I working harder for the good moments (T7.19, T22.11)? Recognize as coercive-control; rebuild outside support.
  • Have I removed the cue (alerts off, app off home screen) and converted reactive checking into scheduled checking with stopping cues restored?

Two or more unresolved boxes on a product: a loop is running against your interest — re-architect the cue, schedule, and friction. Unpredictable warmth-and-coldness from one person over time: treat as a serious red flag and seek trusted outside support.

Key Intel · TakeawaysSummary
  • Conditioning is influence across time. It installs a self-reproducing loop — a habit, compulsion, or attachment — that yields the desired behavior repeatedly without fresh persuasion (Category 18).
  • Two engines. Classical conditioning makes neutral cues inherit the charge of what they predict (Pavlov; Rescorla–Wagner); operant conditioning makes behavior a function of scheduled consequences (Thorndike; Skinner) [C4–C5].
  • Unpredictable reward is the most powerful schedule. Variable-ratio reinforcement produces the highest, most extinction-resistant behavior — the slot-machine schedule engineered into feeds, matches, and loot boxes (Ferster & Skinner) [C5]. You check more when payoff is less predictable.
  • The brain rewards surprise, and wanting ≠ liking. Dopamine encodes reward-prediction error (Schultz); conditioning can sensitize wanting while liking fades (Berridge & Robinson) [C4]. Urge-without-enjoyment is a diagnostic, not a paradox.
  • Habits run on cue → routine → reward. Repetition shifts control to automatic cue-triggered performance (Wood & Rünger); design disciplines (Fogg's B=MAT; Eyal's Hooked) build products around this — neutral tools Category 18 turns against the user [C3–C4].
  • Predatory forms are near-misses and intermittent affection. Near-misses recruit win circuitry despite being losses (Clark et al.); intermittent affection is the conditioning core of trauma bonding — recognized by unpredictable oscillation, countered with outside support [C3]. The defense is structural, not willpower.
Training Op · Red-Team & Operational PracticeAuthorized · Consented
Demonstration ObjectiveLet a class or team feel the difference a reinforcement schedule makes — that an unpredictable payoff drives more checking, persistence, and "just one more" than an identical predictable one — so the slot-machine schedule becomes something participants have watched themselves do. Target "aha": the pull comes from the uncertainty, not from any pending reward; the schedule, not your willpower or the prize, is doing the work.
Exercise Design (consented)Run a consented variable-reward versus fixed-reward demonstration in the schedules-of-reinforcement tradition (after Ferster & Skinner — a demonstration, not unconsented research, with no real gambling, money, or spending of any kind). Use a trivial sanitized tapping task; randomly assign two conditions on the same task — one earns a point on a fixed schedule, the other on a variable-ratio schedule, total payoff held equal. Points are meaningless in-room tokens, not currency. Measure aggregate press rate, persistence after payoff thins, and self-rated urge. Safety rails: consent and immediate debrief; no money or chance-based spending; no minors; results aggregate only.
What Participants CatchThat two people doing the identical task for the identical total reward behave differently purely because of when the payoff arrives; that the variable schedule produces compulsive re-checking and superstitious tap-rituals; that engagement stays high while actual reward is flat (wanting outrunning liking). (S-C-A-R-E-D: R · reward scheduled for persistence · spot-the-loop)
Debrief-to-DefenseReveal the two schedules and name the mechanism: variable-ratio is the extinction-resistant slot-machine schedule engineered into feeds, matches, and loot boxes. Install the structural counters via P-A-U-S-E-D: E (enforce process — convert reactive checking into scheduled checking, add friction) and P (pause). The defense is architectural: attack the cue, break the schedule, restore stopping points, separate wanting from liking. (P-A-U-S-E-D: E · P)
Guardrail & Failure ModesConsent and immediate debrief mandatory; sanitized token task only; no real gambling, money, purchases, or chance-based spend, and no minors used to demonstrate any monetization mechanic; results aggregate and non-identifying. It backfires by introducing real stakes (imports the very harm), using it to build a compulsive product rather than expose one, omitting the debrief, or framing the result as personal weakness (triggers reactance). Absent any safeguard, the demonstration does not run.
Comms Net · Linked SectorsCross-references
  • FM 1-2 · Ch 4 · Ch 5
    Back — cues/salience feed the loop; Zeigarnik pull; loss aversion & emotional cycling
  • Mechanisms §2.4 · §2.22 · §2.23
    Fwd — reward seeking, habit loops, variable reinforcement
  • Mechanisms §2.5 · §2.12 · §2.13
    Fwd — loss aversion, commitment, consistency
  • Category 18 · Category 19
    Fwd — Behavioral Conditioning in full; digital attention optimization
  • T6.1 · T6.9 · T6.13
    Fwd — foot-in-the-door & salami escalation
  • T7.19 · T22.11
    Fwd — intermittent reinforcement & trauma bonding
  • Ch 5 · Ch 7 · Ch 14
    Adjacent — affective rewards; leaderboards/status; habit at technique altitude
§ Doctrinal Sources · SelectedFurther Reading
  • Berridge, K. C., & Robinson, T. E. (1998). "What is the role of dopamine in reward?" Brain Research Reviews, 28, 309–369.
  • Clark, L., Lawrence, A. J., Astley-Jones, F., & Gray, N. (2009). "Gambling near-misses enhance motivation to gamble and recruit win-related brain circuitry." Neuron, 61, 481–490.
  • Duhigg, C. (2012). The Power of Habit. Random House.
  • Eyal, N. (2014). Hooked: How to Build Habit-Forming Products. Portfolio/Penguin.
  • Ferster, C. B., & Skinner, B. F. (1957). Schedules of Reinforcement. Appleton-Century-Crofts.
  • Fogg, B. J. (2009). "A behavior model for persuasive design." Proc. 4th Intl. Conf. on Persuasive Technology, Article 40.
  • Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., & Wardle, J. (2010). "How are habits formed." European Journal of Social Psychology, 40, 998–1009.
  • Pavlov, I. P. (1927). Conditioned Reflexes. Oxford University Press.
  • Rescorla, R. A., & Wagner, A. R. (1972). "A theory of Pavlovian conditioning." In Classical Conditioning II (pp. 64–99). Appleton-Century-Crofts.
  • Robinson, T. E., & Berridge, K. C. (1993). "The neural basis of drug craving: an incentive-sensitization theory." Brain Research Reviews, 18, 247–291.
  • Schüll, N. D. (2012). Addiction by Design: Machine Gambling in Las Vegas. Princeton University Press.
  • Schultz, W., Dayan, P., & Montague, P. R. (1997). "A neural substrate of prediction and reward." Science, 275, 1593–1599.
  • Skinner, B. F. (1938). The Behavior of Organisms. Appleton-Century-Crofts.
  • Skinner, B. F. (1948). "'Superstition' in the pigeon." Journal of Experimental Psychology, 38, 168–172.
  • Thorndike, E. L. (1911). Animal Intelligence: Experimental Studies. Macmillan.
  • Wood, W., & Neal, D. T. (2007). "A new look at habits and the habit–goal interface." Psychological Review, 114, 843–863.
  • Wood, W., & Rünger, D. (2016). "Psychology of habit." Annual Review of Psychology, 67, 289–314.