Data Collection and Graphing
Domain A is 13 of the 75 scored questions (17%) and covers the eight tasks A.1 through A.8 — how you count behavior, how you write it down, how you turn it into a number, and how you read the graph that comes out. It trips people up because the procedures sound nearly identical on paper: partial interval and whole interval differ by one word, and "frequency" and "rate" are used interchangeably in ordinary speech but not on this exam. The 3rd-edition outline also raised the bar — you are now expected to calculate summary measures (A.6), identify trends in graphed data (A.7), and describe the risks of unreliable data and poor procedural fidelity (A.8), none of which the 2nd-edition task list required.
What this domain asks you to do (A.1–A.8)
Read these eight tasks the way the exam writers did: every one of them is something you physically do in a session, not something you know. If you can perform all eight on a real client, on a real data sheet, the domain is yours. Notice the verbs. Six of the eight are "implement," "enter," "calculate," or "identify" — things a behavior technician does. Only A.5 and A.8 are "describe" tasks, and even those are practical: describing a behavior in observable terms is what makes it countable, and describing the risks of bad data is what stops you shrugging off a missed data point. Learn the task numbers as well as the content. Candidates who fail the exam receive a report listing the tasks on which they answered items incorrectly, so knowing that "A.2" means interval recording turns that report into a study plan instead of a mystery.
Continuous measurement: you catch every single instance (A.1)
Continuous measurement means you record every occurrence of the behavior during the observation period — nothing is estimated and nothing is sampled. There are four you must be able to run on demand. Frequency (count) is a simple tally of how many times the behavior happened: five hand-raises, three instances of hitting. Duration is how long the behavior lasted, from onset to offset, either total duration across the session or the duration of each episode. Latency is the time between the moment something is presented and the moment the behavior starts — you start the timer when the instruction, or the trigger, is delivered and stop it when the client begins responding. Interresponse time (IRT) is the gap between the end of one response and the start of the next; it measures the space between behaviors rather than the behaviors themselves. The exam separates these by asking what the clock is doing. Frequency uses no clock at all. Duration times the behavior. Latency times the pause before the behavior. IRT times the pause between two behaviors.
Discontinuous measurement: sampling instead of counting (A.2)
Discontinuous measurement divides the observation into equal intervals and asks a yes/no question about each one. You use it when the behavior is too fast, too continuous, or too hard to time while you are also teaching — vocal stereotypy, on-task behavior, hand-flapping. There are three, and the whole difference is when in the interval the behavior has to occur for you to score it. Partial interval: score the interval if the behavior occurred at any point during it, even for one second. Whole interval: score the interval only if the behavior occurred for the entire interval, start to finish. Momentary time sampling (MTS): ignore the whole interval and look only at the exact moment the interval ends — score it if the behavior is happening at that instant. All three are reported the same way, as a percentage of intervals scored, never as a count. If you observed ten intervals and scored four of them, the answer is 40% of intervals — not "four occurrences." That reporting rule is worth as much on the exam as the definitions.
Which one over-estimates and which one under-estimates
This is the single most tested discrimination in Domain A, and it follows directly from the definitions rather than from memorizing a table. Partial interval over-estimates: a behavior that occupied one second of a ten-second interval earns that interval a full mark, so the resulting percentage makes the behavior look like it filled more of the session than it did. That makes partial interval the sensible choice for behaviors you are trying to reduce, because it will not let you claim success too early. Whole interval under-estimates: a behavior that ran for nine of ten seconds scores a zero, so the percentage makes the behavior look smaller than it was. That makes whole interval the sensible choice for behaviors you are trying to increase, such as on-task or in-seat behavior, because it holds a high bar. Momentary time sampling does not systematically inflate or deflate in one direction; it is a snapshot, so it may miss behaviors entirely or catch a rare one, and it becomes more accurate as the intervals get shorter. Its practical advantage is that you only have to look up at one moment, which is why it is the interval method used when the observer is also running the session or watching a group.
Permanent product recording (A.3)
Permanent product recording measures the lasting outcome a behavior leaves in the environment rather than the behavior itself — you were not there when it happened, but the evidence is. Counting the number of math problems completed on a worksheet, the number of toys left on the floor, the number of dishes washed, or photographs of a bedroom before and after a cleaning routine are all permanent product. Its advantages are real: you do not have to observe in real time, so you can teach without splitting your attention, and a second person can score the same product later, which makes agreement checks easy. Its limitation is equally real: you learn nothing about how the outcome was produced. Twenty completed math problems tell you nothing about whether the client did them, whether a sibling did them, or how many prompts it took. The exam tests exactly that limitation — if the question turns on the process rather than the product, permanent product is the wrong answer, and direct observation is right.
Describing behavior in observable and measurable terms (A.5)
An operational definition describes a behavior so precisely that two people watching the same session, without talking to each other, would record the same thing. It must be observable (you can see or hear it), measurable (you can count it or time it), and clear enough that it says where the behavior starts and where it stops. "Aggression" is not an operational definition. "Closed-hand contact with another person's body from a distance of at least six inches, excluding accidental contact during play" is. The same rule applies to the environment: "the room was chaotic" is an opinion, while "four peers were present and the fire alarm sounded at 10:14" is a description. The test the exam uses is the dead-man test and the stranger test rolled together — if a dead man could do it, it is not behavior ("not talking," "sitting still" as an absence), and if a stranger could not score it without asking you what you meant, it is not operational. Internal states are the classic trap: frustrated, anxious, non-compliant, attention-seeking, wants a break. None of those are observable. What is observable is what the person did.
Entering data and updating graphs (A.4)
Data are entered accurately, completely, and as soon as possible — ideally during or immediately after the session, while you still remember the trial you were unsure about. Reconstructing a session from memory at the end of the week is not data collection; it is guessing, and it corrupts every decision made downstream from it. Record what actually happened, including the parts that make you look bad: the trials you had to prompt more than the plan says, the session that ran short, the target you forgot to run. If you miss a data point, leave it missing and tell your supervisor rather than filling it in with a plausible number. On the graph itself, follow the conventions your supervisor uses: sessions or dates run along the horizontal (x) axis, the measure runs up the vertical (y) axis, each point is one session, and points within the same condition are connected. Points are never connected across a phase-change line — the vertical line that marks where baseline ended and treatment began — because connecting them would imply a continuity between two different conditions that does not exist.
Calculating and summarizing data (A.6)
New in the 3rd edition: you are expected to do the arithmetic, not just record the raw numbers. Rate is count divided by time, and the unit has to be stated — 15 instances of aggression in a 90-minute session is 15 ÷ 90 = 0.17 per minute, or equivalently 10 per hour. Rate is what lets you compare a 30-minute session to a two-hour one; a raw count cannot do that, which is why an exam item that mentions sessions of different lengths is almost always steering you toward rate. Mean duration is total duration divided by the number of episodes: four tantrums totaling 22 minutes is a mean duration of 5.5 minutes. Mean latency is total latency divided by the number of opportunities. Percentage of trials correct is correct responses divided by trials presented, times 100: 12 correct out of 20 trials is 60%. Percentage of intervals is intervals scored divided by intervals observed, times 100. Get in the habit of writing the unit next to every number you produce — "per minute," "per hour," "% of trials," "% of intervals" — because a large share of the wrong options on this exam are the right arithmetic with the wrong unit attached.
Reading a graph: level, trend, and variability (A.7)
Also new in the 3rd edition: you must be able to look at plotted data and say what they show. Three words do almost all the work. Level is where the data sit on the y-axis — high, low, and whether the level changed after a phase change. Trend is the direction the data are moving over time: ascending (going up), descending (going down), or zero/flat. Variability is how much the points bounce around from session to session; highly variable data are unstable and hard to draw conclusions from, which is why baselines are usually continued until they are reasonably stable. Put them together and you can read any single-case graph: "during baseline, aggression was at a high, stable level with no trend; after the phase change, the level dropped and a descending trend emerged." Two cautions the exam likes. First, direction is not the same as good — a descending trend is progress for a behavior you are reducing and a problem for a skill you are teaching, so always check what is on the y-axis before you judge. Second, reading a graph is not the same as deciding what to do about it; noticing a worsening trend is your job, changing the program because of it is your supervisor's.
Two graphs read out loud
Graph one, behavior reduction. The y-axis is instances of aggression per hour, the x-axis is sessions 1 through 11, and a vertical phase-change line sits between sessions 5 and 6. Baseline points are 8, 9, 7, 9, 8. Treatment points are 12, 9, 7, 5, 3, 2. Read it: baseline was at a high level, roughly 8 per hour, stable, with no trend. The first treatment point jumped to 12 — higher than any baseline point — and then the data descended steadily to 2. The correct reading is that the intervention is working: the level dropped and the trend is clearly descending, and that single spike at session 6 is an increase immediately after a procedure was introduced, which is exactly the pattern you would document and report rather than treat as failure. Graph two, skill acquisition. The y-axis is percentage of trials correct, sessions 1 through 10, phase-change line after session 3. Baseline points are 10%, 15%, 10%. Teaching points are 20%, 35%, 30%, 55%, 60%, 75%, 80%. Read it: baseline was at a low, stable level; after teaching began the level rose and the trend is ascending with some variability, since session 6 dipped below session 5. The correct reading is steady acquisition, not a problem — a single dip inside a clear ascending trend is variability, not a reversal.
Unreliable data and poor procedural fidelity (A.8)
The third genuinely new task in this domain asks you to explain what goes wrong when data are unreliable or a procedure is not run as written. Unreliable data means the numbers do not match what actually happened — because the definition was vague, because two people scored it differently, because you filled in a form from memory, or because you scored what you hoped for. Procedural fidelity (also called treatment integrity) means running the intervention exactly as it is written: the same prompt level, the same reinforcer, the same schedule, every session. Poor fidelity produces exactly the same damage as bad data, because the graph now shows the results of a procedure nobody actually ran. The risks are concrete and worth being able to state: your supervisor may conclude a program is not working and discard a procedure that would have worked, or conclude it is working and keep one that is not; the client loses weeks of progress either way; drifting between technicians makes the client's day inconsistent, which itself worsens behavior; and inaccurate records can be a billing and legal problem. The habits that prevent it are equally concrete — use the written definitions, do not modify the procedure to make the session go smoother, and tell your supervisor promptly when fidelity slipped or a definition is not workable.
Key facts and numbers for Domain A
These are the items worth committing to memory before you sit the exam. Everything in this list is either an arithmetic rule or a definition you will be asked to apply to a scenario rather than recite. Practice the calculations with a pen until the unit comes out automatically, because under time pressure the arithmetic is not what fails — the unit is.
Worked scenario: what did you just measure?
An RBT is asked to record Maya's vocal stereotypy during a 20-minute independent-work period. The data sheet is divided into forty 30-second boxes. The RBT is instructed to glance at Maya as each 30-second box ends and mark the box if stereotypy is happening at that instant. At the end of the period, 10 boxes are marked. Which measurement procedure was used and how should the result be reported? Four options are offered: partial-interval recording reported as 25% of intervals; momentary time sampling reported as 25% of intervals; momentary time sampling reported as 10 occurrences; whole-interval recording reported as 25% of intervals. Three of these look right, which is exactly how this exam is built, so eliminate on the details. "Glance as the box ends" fixes the procedure: you are only looking at the moment the interval terminates, so this is momentary time sampling, and both partial interval (any point in the interval) and whole interval (the entire interval) are eliminated on the definition alone. That leaves the two MTS options, and they differ only in the reporting. Interval methods never produce a count, because you did not observe continuously and cannot know how many separate instances occurred — you know only how many of your sampled moments contained the behavior. So "10 occurrences" is wrong, and the answer is momentary time sampling reported as 10 ÷ 40 × 100 = 25% of intervals. Notice the structure: the procedure was decided by one phrase in the stem, and the winner among the survivors was decided by the reporting rule.
Exam traps in Domain A
Every trap below is a pair the exam deliberately puts in the same set of options. Learn them as pairs, not as separate definitions, because on the exam you will never be asked "what is partial interval" — you will be shown a session and asked which of four plausible procedures it was.
Key takeaways
If you remember nothing else from this chapter, remember these. Each one has decided more exam items than any single definition.

Practice stays free. The full RBT (Registered Behavior Technician) study guide is the material itself, taught start to finish — a downloadable PDF + EPUB you keep.