How the 2026 Accounting Firm Digital Front Door Benchmark was conducted
Sample construction, collection ethics, eligibility, calibration, rubric decisions, statistical treatment, and limitations for methodology afdfb-2026-v1.
- Eligible sample
- 118 firms
- Pages observed
- 816 public pages
- Rubric
- rubric.v1.1
Research question and unit
The study asked how well accounting firms' public digital front doors help a prospective client understand fit, establish trust, reach the right service path, and take a clear next step.
The unit was one distinct public website experience for one operating accounting practice. The study measured visible website conditions, not internal business performance.
Sample construction
Candidates came from public OpenStreetMap elements tagged as accounting offices with a public website. The frame was split into four United States and four Canadian geographic strata. Stable SHA-256 ordering and domain deduplication produced a frozen 240-record oversample.
Collection reached 160 accessible sites. After applying the preregistered eligibility rules, 118 firms were eligible: 85 in the United States and 33 in Canada.
Eligibility
Eligible sites visibly offered accounting, bookkeeping, CPA, tax-accounting, client-accounting, or accounting-advisory services and exposed a usable homepage plus at least one relevant public surface or signal.
Directories, associations, educators, regulators, banks, investment-only businesses, software vendors, parked sites, social-only records, duplicate digital experiences, and unidentifiable network pages were excluded with a reason. Inaccessible and robots-restricted candidates were limitations, not zero-scored sites.
Ethical public-site collection
The collector identified itself as public research, validated public destinations before redirects, respected robots.txt, serialized requests per origin, limited global concurrency to two, and used at least a one-second origin delay or the declared crawl delay.
It observed no more than six HTML pages per site and 816 public pages in total. No form was submitted. No call, email, message, chat, booking, upload, authentication, security bypass, or account creation occurred. Full HTML was not retained. Phone numbers and email addresses were redacted from bounded evidence windows.
Rubric and scoring
The calibrated rubric.v1.1 rubric covered nine dimensions: clarity, trust, access, prospect routing, booking, intake and secure handoff, response readiness, human escalation, and mobile or technical foundations.
Supported observations received declared points, unsupported observations received zero, and not-applicable signals were removed from the available maximum. Ambiguous or low-confidence decisions required review and could not silently become failures. The internal composite was not used to rank or identify firms publicly.
Calibration and review
A deterministic 20-site calibration subset received an independent model-assisted review over persisted structural evidence and, when necessary, the public page. This was calibration, not a claim of human inter-rater reliability.
Criteria with excessive disagreement, unresolved cases, or dependence on non-public performance were revised or removed before final scoring. The registered rubric was preserved and the calibrated revision was saved separately.
Two evidence streams
The white paper keeps two evidence streams separate. Structured public-site evidence comes from the 118-site benchmark and is the only source of quantitative prevalence in the report. Practitioner field observations come from the founder's firsthand consulting and commercial experience across more than 100 conversations with CPA and accounting-firm owners.
These were consulting and commercial field observations, not a structured interview sample. The practitioner observations were not formally coded for prevalence, are not results from the 118-site benchmark, and cannot establish causation. They are reported only as recurring patterns and are excluded from the aggregate JSON and CSV downloads.
White-paper qualitative layer
The white-paper edition groups the eligible sites into four mutually exclusive descriptive stages: no clear contact path, contactable only, partly guided, and connected journey. The stages are derived from public contactability plus seven observed signals: an explicit new-client route, explained booking, visible intake, prospect/client separation, a response expectation, after-hours guidance, and a sensitive-document boundary.
This synthesis is not a quality score. A separate manual and model-assisted reading of the existing 20-site calibration subset informed qualitative homepage observations about broad service lists, founder-led identity, generic promises, and limited fit guidance. That review was not reliable enough for prevalence claims. Public mobile-experience quality remains excluded because viewport metadata is an observed technical fact, not a proxy for mobile usability.
Analysis and confidence intervals
Each published proportion came from a named query and retained its numerator, denominator, definition, filters, method version, rubric version, and limitation. Wilson 95% confidence intervals are shown to express sampling uncertainty around the observed sample proportion, not to imply a probability sample.
Subgroup comparisons required at least 20 eligible records per displayed group. Interaction findings required at least 20 observations in every displayed cell and a difference of at least 10 percentage points. These were editorial materiality rules, not statistical-significance tests.
What the study can and cannot establish
The study can describe what was visibly present on the 118 eligible public websites under the rubric during the observation period. It can support discussion of common patterns and design questions within this sample.
It is a geographically stratified convenience sample and is not nationally representative. OpenStreetMap completeness varies, and firms without a qualifying record had no selection opportunity. The study cannot establish national prevalence, lead loss, revenue impact, service quality, internal workflows, response performance, security quality, or causal business outcomes.
Publication-safe aggregate package
The public downloads contain aggregate metrics and the descriptive journey-maturity distribution only. They exclude firm identities, domains, source URLs, captured text, individual decisions, reviewer notes, and firm-level results.