Abstract
Families choosing autism services, clinicians choosing where to practice, and payers authorizing care all share the same reasonable expectation: a provider should be able to explain how it assures the quality of its care and show what that care produces. Outcome reporting in applied behavior analysis has not consistently met that expectation. Reports often lean on satisfaction measures or on the percentage of clients who improved, and they rarely connect results to the clinical systems behind them.
This paper describes how ANNA Autism Care, a center-based provider of Naturalistic Developmental Behavioral Intervention (NDBI) services for autistic children ages 1 through 6 in Massachusetts, builds and assures clinical quality, and reports the outcomes that system has produced for every child in its assessment record. ANNA’s approach rests on three commitments: 1) the clinical model is written down as practice guidelines aligned with national standards, 2) service delivery is verified through treatment fidelity observation, independent treatment plan review, and daily documentation auditing, and 3) measurement is transparent and results are reported as raw-score gains with record counts and medians, so readers can see exactly what each figure represents.
Across 269 comparative assessment measurements from 48 children, 97% of the 96 skill-acquisition comparisons were positive, with a median gain of 21.8 VB-MAPP milestones per reassessment interval. Forty-five of the 46 children with repeated milestone assessments scored higher at the latest assessment than at the first. Adaptive behavior standard scores held or rose against age-advancing norms in about seven of ten comparisons, and measures of interfering behavior improved in two thirds. Dependable outcomes for young children are built on clinical systems that are measurable, transparent, and open to scrutiny. Any provider making outcome claims should be prepared to demonstrate both the integrity of those systems and the results they produce.
Keywords: applied behavior analysis, naturalistic developmental behavioral interventions, outcomes measurement, treatment fidelity, quality assurance, clinical infrastructure, autism, early intervention, center-based services
- 97%Skill-acquisition comparisons positive
- +21.8Median milestones gained per interval
- 45 of 46Clients higher at latest vs. first assessment
- NPS 94Pooled family Net Promoter Score (2026, n = 35)
Section 1
Introduction
The question behind this paper
With autism now identified in approximately 1 in 31 children (Shaw et al., 2025), the demand for applied behavior analysis (ABA) services continues to grow faster than the workforce available to provide it (Behavior Analyst Certification Board, 2026). ABA is supported by decades of evidence documenting gains in communication, adaptive, and social outcomes for autistic children (Smith & Iadarola, 2015), and it is now a covered standard of care across many public and private payers. Naturalistic Developmental Behavioral Intervention (NDBI) is a form of ABA that integrates behavior-analytic procedures with developmental science, teaching skills within play and everyday routines rather than apart from them. As access expands, a central question becomes more important: how does an organization know, and demonstrate, that the care children receive is delivered with quality and produces the outcomes it promises?
Why quality systems matter clinically
Implementation quality is itself a clinical variable. What a child receives in each session depends on how consistently treatment procedures are implemented (DiGennaro Reed et al., 2011; Fiske, 2008), and an organization’s measurement practices determine whether it can evaluate its own effectiveness (Cooper et al., 2020). Outcome reporting requires the same discipline. Satisfaction measures are valuable, but they answer a different question than clinical progress. The percentage of clients who improved shows direction but says nothing about magnitude. Percentage change can also overstate progress on criterion-referenced instruments when baseline scores are low. Stronger reporting pairs the magnitude of change with the number of measurements behind each result and presents outcomes alongside evidence that treatment was delivered as designed. The same logic extends to cost. Programs that produce faster skill acquisition and durable reductions in interfering behavior position children for earlier titration and discharge, and the projected lifetime savings of effective early intervention are substantial (Jacobson et al., 1998; Chasson et al., 2007).
What this paper does
This paper has two aims: 1) to describe how ANNA structures its clinical model and assures the quality of its delivery, and 2) to report client outcomes across the organization’s complete assessment record, with every analytic convention stated in full. ANNA provides center-based services to autistic children ages 1 through 6 across six Massachusetts locations. Families who entrust us with their children’s early intervention deserve to see both the outcomes of care and the clinical system designed to produce them. The same transparency matters to the clinicians and funders who make that care possible.
Section 2
The Clinical Model and How Its Quality Is Assured
A model that is written down
ANNA’s clinical model is documented in the ANNA Clinical Practice Guidelines, which translate professional standards into everyday clinical decisions and align with the BACB Ethics Code for Behavior Analysts (Behavior Analyst Certification Board, 2020), the CASP ABA Practice Guidelines (Council of Autism Service Providers, 2024), and the standards of the Autism Commission on Quality (ACQ). The model is rooted in Naturalistic Developmental Behavioral Interventions (Schreibman et al., 2015) and blends the Early Start Denver Model (Rogers & Dawson, 2010), Project ImPACT (Ingersoll & Dvortcsak, 2019), and Reciprocal Imitation Training (Ingersoll & Schreibman, 2006) around each child’s individual assessment picture.
The guidelines span the full arc of a child’s care and include how medical necessity, service model, and treatment intensity are determined; how each child’s presentation is assessed and conceptualized; and how treatment plans, goals, and protocols are developed from that picture. They set expectations for assent and neurodiversity-affirming practice, active treatment procedures, high-risk procedures, treatment settings, and caregiver coaching, and they define how care is sustained and concluded responsibly through supervision and fidelity, caseload management, cultural and language responsiveness, documentation standards, an appeals process, coordination of care through transition and discharge, and quality assurance and compliance monitoring. The sections are written so that a clinician facing a decision, at any point from intake to discharge, finds written guidance rather than improvising alone.
Assent deserves particular mention, because assent-based practice is central to ANNA’s model. Caregiver consent authorizes treatment; client assent governs whether a specific activity proceeds in the moment, and both are required. The guidelines define assent and its withdrawal in observable terms, specify how staff respond, and require that every child’s program include a goal supporting advocacy, protest, or refusal. Staff are trained to treat the absence of affirmative engagement as a signal to adjust, so that every child has an effective and respected way to communicate no.
Care that is trained, then verified as delivered
A written model matters only if the people delivering it are trained to the model and then verified against it. Every clinician, Behavior Technician (BT) and Board Certified Behavior Analyst (BCBA) alike, completes comprehensive onboarding training in the ANNA model before working independently with children. All training is competency-based and covers all aspects of the ANNA Clinical Practice Guidelines, including NDBI practices, assent standards, safety procedures, and documentation expectations. Each competency must be demonstrated to criterion, with instruction and practice continuing until it is. Verification then carries the same expectations into everyday operations. Behavior Technicians are regularly observed delivering the NDBI practices the model calls for. BCBAs are observed across the full span of their clinical work through a defined set of fidelity checks including a dedicated check on assessment implementation, in which the BCBA is observed conducting the assessment process itself, along with checks covering program oversight, BT supervision, caregiver coaching, and fidelity to the specific NDBI protocols ANNA has adopted. Every treatment plan is reviewed and approved by a Clinical Director before it is submitted. Every clinical session note is audited daily by an automated system against documentation and medical-necessity expectations. Clinical Directors carry a defined set of standing oversight activities, verified monthly with executive attestation.
A system that learns from its own findings
Verification counts for little unless its findings lead to action. ANNA’s Quality Assurance and Improvement program reviews performance quarterly across five domains: treatment fidelity, case management, clinical documentation, client outcomes, and the experience and satisfaction of families and staff. Findings that fall short of expectations become named improvement initiatives with owners, defined actions, and re-measurement dates. Safety incidents are trended within the same review so that clinical and operational signals are read together, and survey results carry defined escalation pathways when they decline. When fidelity observations identify an area of practice that needs strengthening, the response is training and curriculum improvement rather than individual blame. The whole system exists so that the quality of a child’s experience at ANNA never depends on chance. Guiding that work from outside the organization’s own chain of command, ANNA’s Community Advisory Board of autistic self-advocates, family members, and allied professionals advises on how quality is defined and measured, keeping autonomy, communication, felt safety, belonging, and family partnership alongside skill acquisition in the definition of a good outcome.
| System layer | What it means in practice |
|---|---|
| Specify | The clinical model, from medical necessity through discharge, is written as practice guidelines aligned with CASP and ACQ standards. |
| Train | Every clinician completes competency-based onboarding training in the model, demonstrating each competency to criterion before working independently with children. |
| Verify: fidelity | Technicians and BCBAs are regularly observed against written expectations, including fidelity to the named NDBI protocols, with Clinical Director oversight verified monthly. |
| Verify: documentation | Every treatment plan is independently reviewed before submission, and every session note is audited daily by an automated system. |
| Measure | Children are reassessed on a standing cadence with standardized instruments, producing the program-to-date outcomes record reported in Sections 3 and 4. |
| Improve | Performance is reviewed quarterly across five domains; shortfalls become named initiatives with owners and re-measurement dates. |
Section 3
Method
This section describes the population, the measures, and the analytic conventions behind every result that follows. The analysis is retrospective, drawn from ANNA’s complete standardized assessment record as maintained in routine care. There is no comparison condition, so the design is observational and descriptive.
Population
The record includes every current or discharged ABA client with at least one verified standardized assessment between January 2024 and September 2026: 95 children who received center-based services across ANNA’s six Massachusetts clinic sites, all between the ages of 1 and 6 at enrollment. Comprehensive and focused treatment models are both represented, with an average of 21.2 delivered direct ABA therapy hours per week (median 20.3, range 0.7 to 37.1). Intensity is determined child by child through medical necessity rather than a standard package, spanning the comprehensive and focused ranges described in the CASP guidelines (Council of Autism Service Providers, 2024). The 48 clients contributing comparative measurements had a median observation span of 11.6 months between first and latest assessments.
Measures
| Instrument | Type | Desired direction | What it measures |
|---|---|---|---|
| VB-MAPP Milestones | Criterion-referenced | Increase | Verbal behavior and related skills across 170 developmental milestones (Sundberg, 2008) |
| VB-MAPP Barriers | Criterion-referenced | Decrease | Interfering behaviors that impede learning; lower scores are better |
| ESDM Curriculum Checklist | Criterion-referenced | Increase | Developmental skills across ESDM domains (Rogers & Dawson, 2010) |
| Vineland-3 (ABC and subdomains) | Norm-referenced | Increase | Adaptive behavior relative to same-age peers (Sparrow et al., 2016) |
Record definition and analytic conventions
The dataset comprises 799 verified assessment administrations across 95 clients and seven instrument domains, compiled from the source documents in each client’s clinical file. Of 1,212 candidate score entries, 385 were duplicates of the same administration and 28 could not be verified, and the unverifiable entries were excluded and logged, leaving the 799 analyzed here. A comparative measurement is a client’s reassessment in a domain against that client’s immediately preceding administration in the same domain. The record contains 269 comparative measurements from 48 clients. Assessments follow a target 6-month reassessment cadence. Observed intervals tracked close to that target for criterion-referenced instruments (median 5.6 months pooled across the criterion-referenced instruments) and ran near twelve months for the annually normed Vineland-3 (median 11.9 months), and rates are therefore normalized per month between administrations. Four conventions govern all reporting: raw-score change in the instrument’s native units is the primary metric, with percentage change secondary because low starting scores inflate it; medians accompany every mean; movement counts as improved only when it strictly moved in the clinically desired direction, with unchanged scores reported as their own category; and record counts appear with every figure. No client, cohort, or period was excluded; the small number of individual administrations that could not be verified against source records is described in Section 6.
Section 4
Results
- 269Comparative measurements
- 48Clients with a reassessment
- 71%Moved in desired direction
- 9%Unchanged
Children acquired skills consistently on the instruments designed to detect teaching effects
Across the combined criterion-referenced skill-acquisition instruments (VB-MAPP Milestones and the ESDM Curriculum Checklist), 93 of 96 comparisons (97%) were positive. The median gain was +23.8 points per reassessment interval and the median rate was +4.2 points per month in care. On the VB-MAPP Milestones specifically, 88 comparisons across 46 clients produced a median gain of +21.8 milestones per interval over a median interval of 5.7 months. ESDM results point the same way at smaller n. All eight comparisons were positive, with a median gain of +76.5 curriculum points.
97.8% of children with repeated milestone assessments scored higher at their latest assessment than at their first
Viewed longitudinally, the 46 clients with two or more VB-MAPP Milestones administrations gained a median of +37.5 milestones from first to latest assessment over a median span of 11.1 months, and 45 of 46 finished above where they started. The single exception, a six-milestone decline, is reported with the same transparency as the gains and carries the same clinical review as every other signal in the system. Gains concentrated among early learners crossing the VB-MAPP Level 1 to Level 2 transition, where a well-implemented NDBI curriculum should produce its steepest slope. Percentage change on these instruments (median +68% across the skill-acquisition group) is reported for continuity with internal dashboards but overstates typical progress for the arithmetic reasons given in Section 1; the raw milestone counts are the primary basis for interpretation.
Adaptive functioning held or gained ground against age-advancing norms in seven of ten comparisons
The Vineland-3 expresses adaptive functioning relative to same-age peers, and the comparison group advances with the child’s age. A child can therefore acquire real skills and hold a flat standard score because typically developing peers gained at a similar rate. Modest positive movement is the expected signature of genuine progress, and that is what the record shows. Across the composite and adaptive subdomains (143 comparisons from 30 clients), 54% of comparisons rose and 17% held with a median gain of +1.0 standard-score points. Communication and Daily Living Skills were the most consistent movers at 58% and 57% rising. Because the comparison group advances with age, holding or gaining standard scores across a median 11.9-month interval is a meaningful result.
Interfering behavior declined for most children measured, and it remains the focus of active improvement work
Interfering behavior is indexed by the VB-MAPP Barriers, which is scored so that a decrease indicates improvement and which measures behaviors that impede learning, including problem behavior, instructional control difficulties, and prompt dependence. Barriers scores fell in 67% of 30 comparisons with a median reduction of 5.0 points. Most children measured showed fewer behaviors standing between them and the milestone gains above. Fidelity observation independently identified behavior management as the area of practice most in need of strengthening, and an all-staff retraining cycle is underway. The next administrations of this same record will show whether it worked.
Family and staff experience measures complement the clinical record
Across ANNA’s two 2026 family surveys (35 responses pooled), no respondent was a detractor on the shared likelihood-to-recommend item, with survey-level Net Promoter Scores of 95 (Q1) and 93 (Q2) and a pooled score of 94; the Q2 mean was 9.79 of 10. The most requested caregiver training topic was supporting communication. The staff pulse survey’s pooled 2026 eNPS across 91 responses was +43 against a target of +40, with the most recent administration reaching +50. Response rates (roughly a fifth of families and a third of staff per administration) keep these findings directional, and raising participation is itself a tracked measure. Figure 4 summarizes both.
The full record
| Instrument / domain | n | Clients | Improved | Held | Med Δ | Mean Δ | Δ / mo | Med %Δ |
|---|---|---|---|---|---|---|---|---|
| VB-MAPP Milestones | 88 | 46 | 97% | 0% | +21.8 | +23.6 | +3.7 | +61% |
| VB-MAPP Barriers (decrease desired) | 30 | 21 | 67% | 3% | -5.0 | -5.0 | -0.8 | -19% |
| ESDM Curriculum Checklist | 8 | 8 | 100% | 0% | +76.5 | +103.7 | +16.0 | +160% |
| Vineland-3 ABC | 36 | 29 | 50% | 19% | +0.5 | +3.1 | +0.04 | +1% |
| Vineland-3 Communication | 36 | 29 | 58% | 17% | +3 | +5.2 | +0.25 | +5% |
| Vineland-3 Daily Living Skills | 35 | 28 | 57% | 14% | +1 | +3.6 | +0.09 | +2% |
| Vineland-3 Socialization | 36 | 29 | 50% | 17% | +1 | +2.5 | +0.09 | +2% |
| All measurements | 269 | 48 | 71% | 9% | +6.5 | +12.2 | +0.78 | +11% |
Improved and Held use the strict classification of Section 3 (Improved means movement in the clinically desired direction, including reductions on decrease-desired domains); the remainder of each row moved counter to the desired direction. Δ denotes raw-score change between consecutive administrations in the instrument’s native units; Δ / mo is the median change per month; Med %Δ is the median percentage change, reported secondarily.
Section 5
Discussion
The findings support the paper’s central thesis, within the limits of an observational service record. In a clinical system built on documented practice guidelines, protocol-specific fidelity measurement, and audited documentation, the outcome record shows consistent skill acquisition on instruments built to detect teaching effects, with adaptive functioning holding or improving against age-advancing norms. Three specific findings stand out. First, the strongest gains appeared among early learners moving through the VB-MAPP Level 1 to Level 2 range, consistent with the developmental focus of ANNA’s ages 1 through 6 service model. Second, instrument literacy is essential to responsible outcome reporting: criterion-referenced and norm-referenced measures answer different questions, so large raw VB-MAPP gains and modest Vineland standard-score movement are not contradictory. Third, the system also surfaced its less favorable signals, including the behavior-management practice area flagged by fidelity observation, and assigned improvement actions and re-measurement. A credible outcomes program is defined not only by the results it highlights, but by how visibly it responds to the results that need improvement.
The family and staff experience data also add context to the clinical record. Across the 2026 family surveys, no respondent was a detractor on the likelihood-to-recommend item, and the Q2 mean reached 9.79 of 10. Staff experience was similarly favorable, with a pooled 2026 eNPS of +43 against a target of +40 and the most recent administration reaching +50 in August 2026. Because these measures are independent of the standardized assessment data, they offer a different perspective on the system. They also point toward priorities for the coming year. Supporting communication was the caregiver-training topic families requested most often, reinforcing the importance of continued focus on communication-related outcomes.
For the field, the implications are practical. The infrastructure described here is buildable on a moderate scale without research funding. What it requires is: 1) clinical standards specific enough to be trained, observed, audited, and improved, and 2) reporting norms that pair raw-score magnitude, record counts, and medians with the fidelity evidence that makes outcomes interpretable. For payers, third-party verification of exactly this infrastructure is the function of ACQ accreditation, which ANNA is pursuing, and the organization-level records behind every figure in this paper exist and are reviewable. The financial logic follows the clinical one. Programs with stronger quality systems are positioned to produce faster skill acquisition and larger reductions in interfering behavior, and children who acquire skills faster need intensive services for less of their childhood. Quality is therefore among the payer’s most direct levers on lifetime cost, and cost-benefit analyses of effective early intervention have long pointed the same direction (Jacobson et al., 1998; Chasson et al., 2007).
As expectations increasingly include demonstrations of outcomes as well as process quality, providers will need to show both what their clinical system requires and what is observed within it. For payers, verified quality is also the most direct path to value, because programs that teach faster and reduce interfering behavior sooner shorten the arc of intensive treatment and the cost that follows it. For children and families, that transparency is a matter of trust. Quality should be visible, measurable, and open to examination.
Section 6
Limitations and Future Directions
Limitations
Three limitations shape how the results in this report should be read. First, the design is a pre-post single-organization service record with no comparison group. Maturation therefore contributes to criterion-referenced gains in young children, and no causal attribution is claimed. The norm-referenced Vineland-3 partially addresses maturation by construction. Second, the record reflects the documents that could be verified. A small number of score totals and dates could not be confirmed against source records, and those items were excluded and logged for review rather than estimated. The Vineland-3 Maladaptive scales are not yet represented for the same reason. Third, some domains remain thin, with ESDM at n = 8 and Barriers at n = 30, and reassessment intervals differ by instrument (near six months for criterion-referenced instruments and near twelve months for the Vineland-3), so per-interval gains are not directly comparable across instruments; the per-month rates in Table 3 are provided for that reason. These constraints are labeled and addressed rather than hidden.
Future directions
The next iteration of this report is already taking shape, and four commitments will strengthen it. First, successive periods will be pooled into larger cohorts reported with standardized effect sizes. Second, outcomes will be disaggregated by treatment tenure and by measures of intensity and fidelity exposure. This will test whether sustained high fidelity compounds gains. Third, lived-experience outcome measures will be co-developed with the Community Advisory Board. These measures will then be added to the standing assessment battery. Fourth, maintenance of gains will be followed through titration of service hours and discharge.
Statements and Declarations
Conflict of interest
The author is Chief Clinical Officer of ANNA Autism Care, the organization whose clinical program and client data are the subject of this paper. Readers should consider this affiliation when evaluating the reported outcomes; the analytic conventions in Section 3, including strict direction classification and complete-record inclusion, were adopted in part to constrain that risk.
Funding
This work received no external funding.
Data availability
Findings are based on de-identified administrative and clinical records maintained by ANNA Autism Care in the routine provision of services. Given the sensitivity of protected health information, the underlying records are not publicly available.
Ethics and consent
This paper reports a retrospective program evaluation using de-identified service records collected during routine care; no procedures beyond standard care were involved, and no individually identifiable information is reported.
References
Behavior Analyst Certification Board. (2020). Ethics code for behavior analysts.
Behavior Analyst Certification Board. (2026). US employment demand for behavior analysts: 2010–2025.
Chasson, G. S., Harris, G. E., & Neely, W. J. (2007). Cost comparison of early intensive behavioral intervention and special education for children with autism. Journal of Child and Family Studies, 16(3), 401–413.
Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). Applied behavior analysis (3rd ed.). Pearson.
Council of Autism Service Providers. (2024). Applied behavior analysis practice guidelines for the treatment of autism spectrum disorder (3rd ed.).
DiGennaro Reed, F. D., Hyman, S. R., & Hirst, J. M. (2011). Applications of staff training procedures to promote treatment integrity. Behavior Analysis in Practice, 4(2), 4–18.
Fiske, K. E. (2008). Treatment integrity of school-based behavior analytic interventions: A review of the research. Behavior Analysis in Practice, 1(2), 19–25.
Ingersoll, B., & Dvortcsak, A. (2019). Teaching social communication to children with autism and other developmental delays: The Project ImPACT guide to coaching parents (2nd ed.). Guilford Press.
Ingersoll, B., & Schreibman, L. (2006). Teaching reciprocal imitation skills to young children with autism using a naturalistic behavioral approach. Journal of Autism and Developmental Disorders, 36(4), 487–505.
Jacobson, J. W., Mulick, J. A., & Green, G. (1998). Cost-benefit estimates for early intensive behavioral intervention for young children with autism: General model and single state case. Behavioral Interventions, 13(4), 201–226.
Rogers, S. J., & Dawson, G. (2010). Early Start Denver Model for young children with autism: Promoting language, learning, and engagement. Guilford Press.
Schreibman, L., Dawson, G., Stahmer, A. C., Landa, R., Rogers, S. J., McGee, G. G., Kasari, C., Ingersoll, B., Kaiser, A. P., Bruinsma, Y., McNerney, E., Wetherby, A., & Halladay, A. (2015). Naturalistic developmental behavioral interventions: Empirically validated treatments for autism spectrum disorder. Journal of Autism and Developmental Disorders, 45(8), 2411–2428.
Shaw, K. A., Williams, S., Patrick, M. E., Valencia-Prado, M., Durkin, M. S., Howerton, E. M., Ladd-Acosta, C. M., Pas, E. T., Bakian, A. V., Bartholomew, P., Nieves-Muñoz, N., Sidwell, K., Alford, A., Bilder, D. A., DiRienzo, M., Fitzgerald, R. T., Grzybowski, A., Hudson, A., Spivey, M. H., … Maenner, M. J. (2025). Prevalence and early identification of autism spectrum disorder among children aged 4 and 8 years: Autism and Developmental Disabilities Monitoring Network, 16 sites, United States, 2022. MMWR Surveillance Summaries, 74(2), 1–22.
Smith, T., & Iadarola, S. (2015). Evidence base update for autism spectrum disorder. Journal of Clinical Child & Adolescent Psychology, 44(6), 897–922.
Sparrow, S. S., Cicchetti, D. V., & Saulnier, C. A. (2016). Vineland Adaptive Behavior Scales (3rd ed.). Pearson.
Sundberg, M. L. (2008). VB-MAPP: Verbal Behavior Milestones Assessment and Placement Program. AVB Press.
Appendix
Measurement Definitions
| Term | Definition as used in this paper |
|---|---|
| Administration | A single standardized assessment event for one client in one instrument domain. |
| Comparative measurement | A client’s reassessment in a domain, compared to that client’s immediately preceding administration in the same domain. |
| Improved (desired direction) | Strict improvement: an increase on skill-acquisition domains, a decrease on interfering-behavior domains (VB-MAPP Barriers). Zero change is classified separately as Held. |
| Raw change (Δ) | Score at reassessment minus score at the preceding administration, in the instrument’s native units. |
| Rate (Δ / mo) | Raw change divided by months elapsed between the two administrations (30.44-day months). |
| Program-to-date record | All administrations recorded from January 2024 through September 2026, with no cohort, quarter, or client selection. |
