DRO Survived 60% Fidelity. At 40% It Stopped Working for Most Participants.

A parametric evaluation in the Fall 2026 Journal of Applied Behavior Analysis tested differential reinforcement of other behavior at five fidelity levels with commission and omission errors happening together — the combination nobody had tested.

Prior work established that DRO tolerates marginal fidelity errors reasonably well — but it tested errors of commission and errors of omission separately. In practice they arrive together. O'Neill, Rey, Eilers and Craig, publishing in the Fall 2026 Journal of Applied Behavior Analysis, tested the combination.

The finding

Using a reversal design in a human-operant arrangement, they evaluated DRO parametrically at 100%, 80%, 60%, 40% and 20% procedural fidelity, with commission and omission errors implemented together at each level. The result splits cleanly:

  • DRO was efficacious for most participants at 100%, 80% and 60% fidelity.
  • DRO was inefficacious for most participants at 40% and 20%.

The interesting number is 60, not 40. A procedure that still works when four steps in ten are run wrong is more robust than the field's usual fidelity talk implies. But the floor is real and it is not far below.

The limits, in the authors' terms

This was a human-operant arrangement, not children with challenging behavior in a clinic. The authors describe it as a parametric evaluation of fidelity level, and the finding is about where efficacy breaks in that preparation — not a license to run treatment at 60%. Results were characterized as holding “for most participants,” which means individual participants departed from the pattern in both directions. One author disclosed that he sat on the SEAB committee that funded the research, though not as a collaborator at the time of the funding decision.

Why it lands now

CMS's August toolkit pushes states toward documenting that services are properly supervised, and states are writing supervision standards into their manuals — Virginia's new Utilization Workgroup has supervision standards explicitly in its remit. Those instruments count supervision hours. This study is about what supervision is for: it puts a number on the level of implementation degradation at which the intervention stops producing the effect the authorization was granted for. Hours are the thing being regulated. Fidelity is the thing that determines whether the hours did anything.

What you must know or do

  • Clinical directors: if you take fidelity data at all, find out what percentage your DRO programs are actually running at. Most practices collect fidelity checklists and never compute the number, and this study says the number between 60 and 40 is where the answer changes.
  • Supervising BCBAs: score commission and omission errors separately on your next three fidelity checks. The combination is what was tested here, and a checklist that returns a single overall percentage hides which kind of error is accumulating.
  • Anyone reviewing a non-responding case: before changing the procedure, check the fidelity with which the current one is being run. A DRO at 40% and a DRO that doesn't suit the function look identical in the outcome data.
  • Owners: this is the evidence to keep when a reviewer asks why a case ran a given number of supervision hours. Fidelity data tied to an intervention's known breaking point is a stronger answer than an hours log.