Behavioral Experiments in Practice
Behavioral Experiments in Practice
CBT Clinical Practice Simulation

Behavioral Experiments
in Practice

Test the belief. Don’t stage the answer.

A behavioral experiment begins with uncertainty.

The task is not to prove CBT right. It is to create a test that can genuinely teach therapist and client something about a prediction that is maintaining behavior.

You’ll practice
  • Turning broad beliefs into testable predictions
  • Defining evidence before seeing the outcome
  • Manipulating maintaining behavior without abandoning reasonable standards
  • Recognizing behaviors that make the experiment uninterpretable
  • Treating protocol drift as data
  • Processing partial confirmation honestly
  • Separating observation from inference
  • Keeping realistic workplace contingencies visible
Belief
Prediction
Test
Observation
?
Case orientation

Meet Dev

Dev is 35 and works as a senior analyst in a large organization. His reviews are strong. His work is respected.

He is also chronically convinced that his competence is more fragile than other people realize.

Before submitting important work, he rereads, rechecks, rewrites, researches, and delays long after ordinary quality-control requirements have been met.

Then he uses his success as evidence that the extra work was necessary.

Scene 1 · “I know this is irrational”
Dev

“I know the whole ‘imposter syndrome’ thing.”

“I know objectively I’m probably not terrible at my job.”

“But if I stop being this careful, that’s when people are going to find out.”

What belief would you test?

Clinical Support

Test the prediction that maintains the behavior.

Global belief
“I’m not actually competent.”
Conditional assumption
“If I stop overchecking, my incompetence will become visible.”
Testable prediction
“If I submit after one structured review rather than repeated checking, there will be an important error and my manager will question my competence.”
“What do you expect will happen if you behave differently?”
“What would another person actually do or say?”
“What part of the belief could we observe?”
Scene 2 · Define the evidence
Dev

“Okay, so I send something without checking it obsessively.”

“Then if nothing happens, we say the belief was wrong.”

What needs clarification before the experiment?

Clinical Support

Decide what the data mean before you know the result.

Supports prediction
Consequential error + explicit questioning of Dev's competence or reliability.
Weakens prediction
No consequential error, or an ordinary correction without the predicted competence judgment.
Ambiguous
No feedback, unclear causal relationship, or a task too trivial to activate the belief.
A useful behavioral experiment does not need to guarantee disconfirmation. It needs to create interpretable information.
Possible experiments

Which task would actually test the prediction?

Task A

Internal email

Normal checking: ~10 minutes

Predicted fear: 30/100

Dev already believes the stakes are low.

Task B

Routine monthly analysis

Normal checking: 60–90 minutes

Predicted fear: 65/100

Meaningful enough to activate the belief.

Task C

Executive report

Normal checking: several hours

Predicted fear: 90/100

Professional consequences of genuine error are high.

Dev

“The email is obviously the one I want.”

Decision · Choose the experiment
Clinical Support

Manipulate the maintaining behavior—not professional responsibility.

Reasonable professional behavior

Standard checklist

Required automated checks

Normal peer review where workflow requires it

Correction of identified factual errors

Potential maintaining behavior

Repeated rereading

Checking the same formula again

Research that does not change the conclusion

Delaying solely to reduce doubt

Scene 4 · “Can my colleague look at it?”

Dev agrees to complete the monthly analysis, use the standard checklist once, correct identified errors, and submit.

Dev

“Can I have my colleague look at it before I send it?”

How do you respond?

Clinical Support

Ask what will get credit if things go well.

“If the report turns out fine, what will you think made it fine?”
“What behavior would let the old belief keep all the credit?”
“Is this ordinary workflow—or added protection because of the experiment?”
Safety behavior is defined by function, not by the surface appearance of the behavior.
Behavioral experiment plan
Belief
“If I stop overchecking, people will discover I’m not as competent as they think.”
Prediction strength
75%
Prediction
Less repeated checking → meaningful error → manager notices → competence questioned.
Manipulation
One structured quality review rather than repeated checking.
Normal supports retained
Standard template · required automated checks · source data access.
Behaviors reduced
Repeated rereading · repeated formula checking · reassurance review · unnecessary research.
Observable data
Errors · severity · manager feedback · correction requests · urge to recheck · time spent.
First attempt · Protocol drift

The experiment changes while Dev is doing it.

Structured check complete
“I could make this sentence clearer.”
Rewrite sentence
Reread paragraph
Check adjacent table
Check another formula
25 minutes pass
Decision · What does the drift mean?
Clinical Support

Protocol drift is data.

Uncertainty
“One more improvement”
Check
Short-term relief
New doubt
Check again
“Where did the plan change?”
“What were you predicting at that moment?”
“What did the extra checking accomplish immediately?”
Revised experiment
Complete normal structured review
Correct identified errors
Review complete
Submit within five minutes
Five minutes is not inherently therapeutic. It targets the period in which Dev's redundant checking expands.
One month later

The experiment happens.

Dev completes the monthly analysis.

He performs one structured review.

He finds two errors and corrects them.

He reaches the end of the checklist.

Urge to check again
8 / 10

Prediction: “There is probably still something important wrong.”

Later that afternoon
Manager message
“Thanks. Can you update the label on Figure 3? It still says Q2 instead of Q3.”
There was an error.
Decision · Inconvenient data

What do you do with the error?

Clinical Support

Do not protect CBT from inconvenient data.

“Which part of the feared outcome happened?”
“Which part did not?”
“Was the magnitude what you predicted?”
“What did the outcome mean?”
“What would you predict now?”
Prediction comparison
Predicted
Less checking
Important error
Manager notices
Competence questioned
Observed
Less checking
Label error
Manager notices
Correction requested
Correction accepted
Partial match ≠ full confirmation.
Post-experiment processing
Dev

“So I was right.”

You ask: “About which part?”

Dev

“That there’d be a mistake.”

“But she didn’t act like it was some huge thing.”

After Dev corrected the label, his manager replied:

“Perfect, thanks.”
Dev

“Although maybe she was annoyed and just didn’t say it.”

Decision · Observation or inference?
Clinical Support

Behavioral experiments produce evidence—not omniscience.

Observed

A labeling error occurred.

Correction was requested.

Dev corrected it.

Revision was accepted.

No negative performance feedback occurred.

Possible but unobserved

The manager may have felt annoyed.

The manager may have privately noticed the lapse.

The manager may have formed some judgment.

Not established: that the error changed the manager's view of Dev's competence.
Belief update
Original confidence
75%
Current confidence
50%
Dev

“I still think checking catches mistakes.”

“I’m less convinced that every mistake turns into some judgment about whether I belong here.”

“But one report isn’t enough.”

A lower belief rating is not automatically a better experiment.
Decision · What now?

What is the best conclusion?

Clinical Support

The outcome of an experiment is a formulation update.

“What changed?”
“What part survived?”
“What new prediction appeared?”
“Did we test the mechanism we thought we were testing?”
“What alternative explanation remains?”
“What experiment would discriminate between those explanations?”
Experiment 2 · A new rule emerges
Dev

“People at my level are supposed to know the answer.”

Guess
Answer vaguely
Promise to double-check
Overexplain

Proposed behavior:

“I’m not sure. I’d need to check that.”
Dev

“That sounds worse than the report.”

Decision · What are you testing?
Cross-team meeting
Colleague

“Do we know whether that change affected the regional numbers too?”

Dev

“I’m not sure. I’d have to check the regional cut separately.”

Colleague

“Okay. Could you send that after?”

The discussion moves forward.

Senior colleague · ten minutes later

“Actually, Dev, can you walk us through the assumption behind your main model?”

Dev does. The conversation continues normally.

Anxiety afterward
6 / 10
Urge to explain himself further
Moderate
Dev

“That was deeply unpleasant.”

“And apparently very boring to everyone else.”

Decision · Consolidate the learning
Behavioral Experiment Learning Record
Experiment 1

Reduce repeated checking

Prediction: error → competence judgment

Outcome: minor error → routine correction

Learning: ordinary errors remain possible; the predicted professional consequence was not observed.

Experiment 2

Allow visible uncertainty

Prediction: “I don’t know” → credibility loss

Outcome: follow-up requested; expertise subsequently sought

Learning: visible uncertainty did not remove Dev from a position of competence.

Not established: nobody will ever judge Dev, mistakes never matter, or checking is useless.
Decision · Realistic threat

How do you respond?

Clinical Support

Realistic threat changes experiment design.

Observed context
+
Client prediction
+
Behavioral cost
+
Ethical risk
Experiment design
“What scrutiny has the client actually observed?”
“Who appears to receive latitude here?”
“What protection is proportionate to the environment?”
“What behavior exceeds what the context requires?”
“What can ethically be tested?”
Integration · Ten minutes remain
Dev

“I think I’ve been using overwork to keep myself from ever finding out how much competence I actually have without it.”

“But I also don’t want to swing to being careless.”

What would you do next?

Pattern synthesis

Your choices shaped what counted as evidence.

This is not a score. It describes what your decisions tended to emphasize in this simulation.

Completed behavioral formulation
Important task / visible uncertainty
“I may not actually be competent.”
“If I don’t overprepare, people will find out.”
Check / overprepare / overexplain / hide uncertainty
Short-term certainty
Good outcome credited to overpreparation
Belief remains untested
Alternative learning path:
Reduce excess protection → observe what actually happens → update formulation → design next test.
Reusable behavioral experiment template

Build the test before you know the answer.

1. Belief
What belief is maintaining the behavior?
2. Conditional prediction
If ______ then I predict ______.
3. Confidence
____ %
4. Behavior to manipulate
What specifically changes?
5. Normal supports
What remains because it is reasonable?
6. Protective behaviors
What might prevent interpretation?
7. Supports prediction
What would count?
8. Weakens prediction
What would count?
9. Ambiguous
What would tell us little?
10. Outcome
What happened?
11. Observed
What was directly seen or heard?
12. Inferred
What are we adding?
13. Unknown
What cannot be concluded?
14. Updated prediction
What changed?
15. Next test
What would teach us something new?
Intervention comparison
Thought record

What evidence supports or challenges this interpretation?

Behavioral experiment

What happens if we change something and observe the result?

Exposure

What can be learned by approaching a feared condition while reducing avoidance or protection?

Problem-solving

Given a real practical problem, what action is workable?

These interventions can overlap without becoming interchangeable.
Debrief

Reflect on your clinical choices.

When did Dev’s belief become experimentally testable?
Why wasn’t the lowest-risk task necessarily the strongest experiment?
What would have made colleague review interfere with Experiment 1?
What did protocol drift reveal?
Why did the labeling error matter?
Which part of Dev’s prediction was confirmed?
Which part was not?
Where could the therapist have made CBT “win”?
What remained unknowable?
When did Experiment 2 resemble exposure?
How should realistic workplace scrutiny affect experiment design?
What experiment would you design next?
Continue the learning

See CBT move from explanation
to experiment.

Behavioral experiments make CBT genuinely empirical when they are designed to test predictions rather than confirm conclusions the therapist has already decided are true.

For longer-form examples of CBT formulation, questioning, uncertainty, behavioral possibilities, and revision unfolding in live conversations, continue with Transcripts, Notes and Reflections from The CBT Dive Video Podcast.

Explore The CBT Dive book