Understanding CBM
Curriculum-Based Measurement is how schools answer a question that a single test score cannot: is this student improving enough in response to what we are actually doing? This is a working reference on how CBM measures are built, scored, graphed, and most importantly, driven by fidelity.
CBM uses brief, repeatable measures of important academic skills. The measures are short by design, consistent in how they are given and scored, and built to be sensitive to relatively short-term change in academic performance. That combination is what makes them useful: a teacher can collect them without disrupting instruction, and a team can graph the results providing the feedback needed to drive intevention adjustments. Consistency is also what makes one probe comparable to the next — and standardized commercial systems add a further layer on top of it: formal technical validation, norms, benchmarks, and alternate-form research.
The National Center on Intensive Intervention identifies frequent progress monitoring as a core component of Multi-Tiered Systems of Support and Data-Based Individualization. Data are collected and graphed during intervention, then evaluated against a student goal to determine whether progress is sufficient.
Why repeated data matter
A single academic score describes performance at one moment. Repeated measurement describes a trajectory, and trajectory is what instructional decisions actually depend on.
Consider two students with the same current score. One started well below expectation and has climbed steadily. The other has been essentially flat for six weeks. Their scores today are identical. Almost nothing else about them is.
This distinction sits at the center of evidence-based intervention. A student may remain below grade-level expectation while demonstrating strong growth — evidence of a positive response and movement toward the instructional goal. Another student may sit at a similar current score with no measurable improvement. Those patterns can lead to very different instructional decisions, and only repeated measurement tells them apart. The gap itself narrows only when a student’s rate of growth exceeds the rate needed to maintain it.
The research base reflects this use. CBM has been examined for screening, benchmarking, progress monitoring, goal setting, intervention evaluation, and instructional decision making. Oral-reading CBM has shown a strong association with standardized reading achievement in the elementary grades, while contemporary reviews continue to document a substantial evidence base for mathematics CBM.
Baseline, benchmark, norm
These three words are used interchangeably in meetings all the time, and they mean entirely different things. Getting them confused is the most common way CBM data get misread.
- Baseline
- The student’s starting level of performance. Teams may establish it from a recent appropriate screening score, or from several initial comparable probes rather than one potentially unrepresentative score. A common approach is to administer three probes and use the median.
- Benchmark
- An established performance target belonging to a specific assessment system. Benchmarks are usually criterion-referenced and may indicate the likelihood of a later outcome.
- Norm
- How a student’s performance compares with a reference population, commonly reported as a percentile rank.
The University of Oregon publishes DIBELS 8 administration standards, benchmark goals, national percentile information, and supporting technical research, including updated percentile tables and revised Zones of Growth. Those publications are a good illustration of the standardization work that has to happen before a score can legitimately be read against external expectations.
Oral reading fluency
Oral Reading Fluency measures performance while a student reads connected text aloud, often as a brief timed passage — commonly one minute, depending on the assessment system. The primary score is Words Correct Per Minute, considered alongside accuracy and the pattern of errors.
ORF is efficient because fluent connected-text reading requires several underlying processes to operate together — decoding, word recognition, and language processing all have to be working for the score to be high. For students with reading difficulty, repeated ORF measurement is a sensitive indicator of change over time.
It is not a complete assessment of reading. Although oral reading fluency is associated with broader reading achievement, vocabulary, language comprehension, inferential reasoning, and extended-text comprehension are not directly or comprehensively assessed by a rate-and-accuracy measure. ORF should be read as one indicator inside a broader reading assessment system. When standardized norms or benchmarks are needed, use a technically validated system and follow its administration and scoring procedures — DIBELS 8, Acadience Reading, aimswebPlus, and FastBridge are established examples.
MAZE
MAZE is a brief silent-reading measure. Selected words in connected text are replaced with a set of choices, and the student selects the word that preserves the meaning of the passage. Administration is silent and can be done with a whole class at once, which makes it practical at the secondary level.
Standardized MAZE measures are designed to capture aspects of comprehension and meaning construction. Performing well requires the reader to combine word recognition, linguistic knowledge, background knowledge, context, and reasoning to work out which word fits. Acadience Reading 7–8, for example, identifies Maze as a measure of reading comprehension; its standardized procedure replaces approximately every seventh word, uses three-minute passages, records correct and incorrect selections, and calculates an adjusted score that accounts for incorrect responses.
One detail from that manual makes the point of this whole page concrete. Acadience 7–8 benchmark assessment administers three Maze passages as a triad and combines them within its own scoring system. Those benchmark interpretations therefore cannot be applied to a single locally generated Maze passage — not even one that follows the same replacement interval and the same three-minute timing.
Like ORF, MAZE is one indicator rather than a comprehensive assessment. Inference, text structure, extended-text understanding, and language comprehension may all require additional assessment.
Mathematics CBM
Mathematics CBM applies the same logic to defined mathematical skills. Measures may examine computation, automaticity, concepts and applications, or another specified target behavior. Depending on the system, scores may be based on correct digits, correct responses, accuracy, rate, or another precisely defined metric — and those are not interchangeable. Correct-digit scoring in particular gives credit for partially correct work, which can capture changes in computation performance that whole-item scoring does not.
The research base extends from early mathematics through the secondary grades, though the amount and type of evidence differ by measure and age. A 2023 review identified 99 studies examining mathematics CBM from preschool through grade 12. Most examined measurement and screening properties; considerably fewer looked at repeated progress monitoring, and only five addressed instructional utility.
That gap is worth sitting with. Collecting a score is not the same thing as using the score well, and the second is where the benefit actually comes from.
From scores to decisions
CBM becomes useful when repeated measurement is placed inside a structured decision process. A strong process has a clearly defined target skill, a baseline, an appropriate goal, repeated comparable measures, consistent administration and scoring, a graph, and rules for reviewing the data that were decided before the data came in.
The National Center on Intensive Intervention describes several approaches for analyzing academic progress-monitoring data, including four-point analysis, trend-line analysis, and analysis of the median of recent observations. All of them exist for the same reason: to move teams away from reacting to one high or low score and toward systematic interpretation of a pattern.
Fidelity changes the meaning
Student outcome data cannot be interpreted without knowing what was actually delivered. If an intervention was designed for five sessions a week and happened twice, limited growth is not evidence that the intervention does not work. When fidelity is inadequate, a team cannot confidently attribute poor growth to the intervention itself.
This is why strong MTSS systems track both student response and implementation fidelity. NASP’s Domain 1 expectations specifically address using systematic, reliable, and valid data to monitor intervention response, interpreting universal screening and progress-monitoring data, and incorporating treatment-fidelity data when decisions are made about modifying or ending an intervention.
Progress monitoring lets a team ask two separate questions: is the student responding, and was the intervention implemented as intended? Both matter, and only one of them is visible on the graph.
CBM within RTI and MTSS
Progress monitoring is fundamental to the logic of Response to Intervention, Multi-Tiered Systems of Support, and intensive intervention. Students receive instruction matched to an identified need. Performance is measured repeatedly. The team evaluates response and decides whether to continue, adjust, intensify, or look more closely at the problem.
Arizona Department of Education guidance describes RTI/MTSS in the same terms: identify instructional needs, provide targeted research-based interventions, use progress monitoring to measure student response, and evaluate whether those interventions were effective.
CBM does not replace professional judgment. It improves professional judgment, by giving a team a systematic record of what happened after an instructional decision was made.
Evaluation and eligibility
IDEA does not require a particular CBM product, and does not require CBM as the sole method of evaluating intervention response.
IDEA does require that teams considering Specific Learning Disability consider data-based documentation of repeated assessments of achievement at reasonable intervals, reflecting formal assessment of student progress during instruction. Where an RTI process is used, that documentation also includes the instructional strategies used and the student-centered data collected.
The regulation does not name a measurement model. OSEP has clarified that §300.309 does not use the term “continuous progress monitoring” and does not mandate one particular approach — what it requires is consideration of repeated achievement data, not any specific product or system.
Progress-monitoring data can therefore provide important evidence about instructional response. No single CBM score determines disability or eligibility, and evaluation decisions should integrate multiple relevant sources of information.
Matching interpretation to measure
There are two legitimate but different uses of CBM-style measures, and the difference is not about quality. It is about what the resulting number is licensed to mean.
A standardized measure is administered, scored, and interpreted according to a validated assessment system. When the required procedures are followed, the system may provide national norms, benchmark goals, risk classifications, and expected growth information. Those interpretations are the product of standardization, norming, reliability and validity research, alternate-form work, and growth-sensitivity studies.
A locally generated probe is still genuinely useful for formative progress monitoring. A teacher can establish a student’s own baseline, set a goal, and measure repeatedly using unfamiliar probes of reasonably comparable difficulty. That can provide useful formative evidence of change, when comparable unfamiliar forms are administered and scored consistently. What it does not produce is a percentile.
The interpretation has to match the measure. External percentile ranks and standardized benchmark classifications should not be assigned to a locally generated probe unless validation supports it. A DIBELS benchmark belongs to DIBELS passages administered under DIBELS conditions. An Acadience Maze benchmark belongs to Acadience Maze — not to any passage that happens to replace every seventh word.
This is not a limitation specific to any one tool. It is a basic measurement principle, and it cuts both ways: changing the passage, timing, scoring rule, or administration conditions of a standardized assessment may limit or invalidate the standardized interpretation of that score.
Keeping it practical
Progress monitoring only helps when it actually gets done. CBM measures are brief by design, and the assessment burden is deliberately kept low so that the monitoring actually happens. Administration can happen on printed probes and scoring forms, or digitally with automated scoring. Both can support efficient monitoring; when using a standardized measure, follow the developer's required administration conditions.
The purpose is not to add another assessment requirement. It is to gather the smallest amount of useful data needed to make better instructional decisions: recognize improvement earlier, identify inadequate response sooner, evaluate whether the intervention was delivered, communicate progress more clearly with families, and change instruction based on evidence rather than impression.
That is why CBM has stayed at the center of academic progress monitoring for four decades.
FarPoint’s free CBM tools
FarPoint provides free browser-based tools for creating, administering, scoring, and printing curriculum-based measures. Nothing is stored and no student data leaves the browser — scores are produced for you to record wherever you already track them.
These tools generate locally created probes. They are built for formative progress monitoring against a student’s own baseline and goal, and they are not a substitute for a standardized assessment when normative interpretation is required. For teams that want a structured longitudinal record, REDAssist organizes repeated CBM scores alongside baselines, goals, aimlines, trendlines, intervention phases, and fidelity data.
The tool is secondary to the principle: measure what matters, measure it consistently, look at growth across time, verify the intervention was actually delivered, then decide what happens next.
References
Primary sources for the standards and requirements described above.
Federal regulation
- Determining the existence of a specific learning disability, 34 C.F.R. § 300.309 (2017). §300.309(b)(2) requires consideration of data-based documentation of repeated assessments of achievement at reasonable intervals, reflecting formal assessment of student progress during instruction.
Professional standards
- National Association of School Psychologists. (2020). The professional standards of the National Association of School Psychologists. Domain 1, Data-Based Decision Making: multiple data sources, progress and outcome measurement, intervention evaluation, and decision making within MTSS.
National technical assistance
- National Center on Intensive Intervention. (2024). Academic progress monitoring tools chart. Technical adequacy criteria: reliability, validity, sensitivity to growth, alternate forms, administration and scoring, rates of improvement, and end-of-year benchmarks. Inclusion on the chart is not endorsement.
- National Center on Intensive Intervention. (2024, December). Decision rules for analyzing academic progress monitoring data. Source for the four-point rule, trend-line analysis, and the median-of-recent-observations approach.
Standardized measurement systems
- University of Oregon. (2026). DIBELS 8th Edition 2024–25 benchmark percentile ranks (DIBELS Technical Report No. 26-002). Current percentile tables through grade 8, with guidance on appropriate interpretation of percentile scores and reference groups.
- University of Oregon. (2026). DIBELS 8th Edition benchmark assessments revised Zones of Growth based on the 2024–2025 school year (DIBELS Technical Report No. 26-003). Revised growth expectations derived from the 2024–25 data.
- University of Oregon. (n.d.). DIBELS 8th Edition materials. Official entry point for the Administration and Scoring Guide, benchmark materials, progress-monitoring probes, and benchmark goals.
- Abbott, M., Good, R. H., III, Gray, J. S., Warnock, A. N., & Powell-Smith, K. A. (2020). Acadience Reading 7–8 assessment manual. Acadience Learning. Identifies Maze as a standardized measure of reading comprehension; documents replacement of approximately every seventh word, three-minute administration, and scoring as correct responses minus one-half incorrect.
State guidance
- Arizona Department of Education, Exceptional Student Services. (2024). AZ-TAS evaluation process. Describes RTI/MTSS as identifying instructional needs, providing targeted research-based interventions, and using progress monitoring to measure response and verify intervention effectiveness.
- Arizona Department of Education. (2025). Universal literacy and dyslexia screener guidance: K–3 multi-tiered system of support literacy assessment and instruction. States that effective MTSS requires robust progress monitoring, and that those data should inform intervention intensity, duration, frequency, and change decisions.
Peer-reviewed research
- Reschly, A. L., Busch, T. W., Betts, J., Deno, S. L., & Long, J. D. (2009). Curriculum-based measurement oral reading as an indicator of reading achievement: A meta-analysis of the correlational evidence. Journal of School Psychology, 47(6), 427–469. https://doi.org/10.1016/j.jsp.2009.07.001 Meta-analysis demonstrating a strong association between CBM oral-reading performance and standardized reading achievement in grades 1–6.
- Nelson, G., Kiss, A. J., Codding, R. S., McKevett, N. M., Schmitt, J. F., Park, S., Romero, M. E., & Hwang, J. (2023). Review of curriculum-based measurement in mathematics: An update and extension of the literature. Journal of School Psychology, 97, 1–42. https://doi.org/10.1016/j.jsp.2022.12.001 Review of 99 mathematics CBM studies from preschool through grade 12. Most addressed measurement and screening properties; only five addressed instructional utility.