Skip to main content

Testing Our Proposed FTA Safety-Data Research Method on FTA's Public Data

Before submitting a concept paper for FTA research topic FTA-SDA-001, we ran parts of our proposed method against FTA's published safety data. This note records what we did, what happened, and what the results do and do not show.

Ben Greene · Co-founder, Transit Automations

Summary. We are proposing research to help the Federal Transit Administration document its safety-data collection processes, define and measure data quality, and keep that work accurate as reporting requirements change. Before asking FTA to fund it, we wanted evidence that parts of the method can do useful, repeatable work on FTA's real public data, not just on paper. FTA's published CY2025 national major-event fatality total was 337 (May 2026 release). Using FTA's public event-level dataset and FTA's documented definitions, the frozen computation also produced 337, with exact agreement across all 13 person-type values and all 21 workbook mode rows. Three data checks also ran over the full public dataset, each against an explicit denominator. A change-replay protocol ran end to end as well. Its misses, each traced to a cause, show the limits of working from change notices and manuals alone. That is worth knowing before the funded study is designed.

Why we ran the demonstration

Transit Automations is preparing a concept paper for FTA's FY2026 research announcement, under topic FTA-SDA-001, Safety Data Management Strategy. The proposed research would help FTA document how its safety data is collected, catalog the systems and processes it uses, and write data-quality rules that trace to specific reporting requirements. It would also test whether that work stays accurate when requirements change.

FTA asks applicants to state how mature their approach is and to back that statement with analytic and empirical evidence. We wanted the evidence to come from running parts of the method. A description alone would not do. So before submitting, we ran a small, bounded demonstration on FTA's own public safety data, under rules we fixed before looking at any result.

This note is the public record of that work. It is neither a sales document nor a review of FTA's data. It reports what we did, what happened, and how far the results reach.

What we set out to demonstrate

The question was narrow: can critical parts of the proposed method do useful, repeatable work on real FTA public data, rather than existing only as a paper concept?

We tested four things:

  1. Checks. Turn documented FTA statements into checks that run over FTA's public event data, each with stated applicability and an explicit denominator.
  2. Reproduction. Trace one FTA-published safety figure back to FTA's event-level data, using FTA's documented definitions.
  3. Change replay. Put one documented FTA reporting change through a replay protocol fixed in advance, and account for everything it found and missed.
  4. Reproducibility. Keep the inputs, rules, and records needed to rerun the work and get the same results.

The demonstration was deliberately small. It was not a production system or an integration test. Nor was it an evaluation of how well FTA or any transit agency reports.

What data and documents we used

Every input is a public FTA or Federal Register product, and a U.S. Government work. We retrieved them on October 6, 2026, recorded a SHA-256 fingerprint of each file, and froze them before any check ran.

  • Major Safety and Security Events (FTA, on data.transportation.gov). The complete public export: 115,454 event records from January 2014 through May 2026, as of FTA's September 9, 2026 update, together with the dataset's published field descriptions.
  • Safety & Security Major Event Time Series (Major-Only) workbook, May 2026 release, in a file dated September 1, 2026. This is the source of the published figure. FTA has since posted a June 2026 release; this demonstration used the May 2026 file, which remains available at the link under Sources.
  • 2025 NTD Safety and Security Reporting Policy Manual, Version 1-1 (November 2025). It provided the reportability criteria and the pre-change baseline for the replay.
  • 2026 NTD Safety and Security Reporting Policy Manual, V1, which documents the Report Year 2026 changes.
  • Federal Register notice 90 FR 30771 (July 10, 2025), National Transit Database Reporting Changes and Clarifications for Report Years 2025 and 2026. The replayed change comes from this notice.

We used no FTA internal data and no data from any transit agency or partner.

What we did

Validate: three checks and a published-figure reproduction. The demonstration used three checks, each tied to a cited FTA requirement, definition, or documented statement:

  • Reportability evidence for calendar-year 2025 non-rail collisions. FTA's 2025 manual says a reportable non-rail collision is one that results in an injury requiring transport away from the scene, a fatality, an evacuation for life-safety reasons, property damage of $25,000 or more, or towing of the transit or non-transit vehicle. The check asks whether each applicable published record carries field evidence of at least one of those criteria.
  • Fatality-total consistency, 2016 onward. FTA's workbook defines the fatality total as the sum of specific person-type categories. The check asks whether each record's published total equals the sum of its 13 person-type columns.
  • Incident Number uniqueness, all years. The dataset's field description calls Incident Number "a unique system-generated identification number for each Major Safety Event." The check asks whether that holds in the public file.

Every check reports four explicit states: pass, fail, not applicable, and unknown (meaning the published fields cannot decide it). Each count is reported against a stated denominator.

For the reproduction, we took the figure FTA publishes in its May 2026 workbook for calendar-year 2025 national major-event fatalities. We then recomputed it from the event-level dataset using FTA's documented definitions: select the 2025 events, apply the workbook's documented exclusions, and add up the 13 person-type fatality columns. The trespasser column is left out, because the workbook documents it as a subtotal of people already counted in other columns. We also compared every person-type value, and every mode's fatality and event counts, with the workbook's own breakdown. A total can agree by coincidence while its parts disagree. The breakdowns test the parts.

Before anything ran on the full dataset, the checks, the reproduction, and the replay procedure were tested on 580 labelled fixtures. These are small constructed cases, some of them deliberately altered copies, each with its expected result written down in advance. All 580 passed. Fixtures test mechanics only; nothing found in a fixture is a finding about FTA data.

Maintain: a change-replay protocol. The proposed research also asks whether a catalog of reporting elements, and the rules built on it, can keep up when FTA changes a requirement. This demonstration could not answer that question. An informative test needs a change chosen before the catalog is built, and we already knew FTA's documents for any change we might have picked. What we could test was the control protocol such a study would use.

We replayed one documented change. The July 2025 notice adds "disabling damage" as a subset of "substantial damage" for rail collision events. The protocol fixes each step before the next one starts, and each step was committed to version control in this order:

  1. Change input. The change terms, taken from the notice's section on this change and from nothing else.
  2. Catalog. 52 entries transcribed from the Collisions subsection of the 2025 manual, each machine-checked against the manual's text.
  3. Reference set. The catalog entries and checks that should need review, and any new element, prepared from FTA's documents. Transit Automations (Ben Greene) adopted the set before the prediction ran, and only its fingerprint was recorded at that point. It is our reference, not an FTA answer key.
  4. Prediction. A deterministic procedure looks for the change terms in each catalog entry's text and follows recorded field links from those entries to our checks.
  5. Reveal and comparison. The reference set was revealed only after the prediction was committed. Every miss and every spurious flag was then tagged with a cause from a fixed list written before the run.

The edition we used as the pre-change baseline, Version 1-1 of the 2025 manual (November 2025), was issued after the July 2025 notice. Its Collisions subsection shows no trace of the change. FTA does not host the pre-notice 2025 edition, so we could not compare against it.

What happened

FTA's published CY2025 national major-event fatality total was 337 (May 2026 release). Using FTA's public event-level dataset and FTA's documented definitions, the frozen computation also produced 337, with exact agreement across all 13 person-type values and all 21 workbook mode rows. The documented exclusions removed nothing from the 2025 data. Every alternative calculation we had declared in advance also gave 337.

The three checks ran over all 115,454 records, and none produced an undecidable result. For calendar-year 2025 non-rail collisions, of 6,649 applicable records, one carried no public-field evidence of the documented criterion that made the collision reportable. That record may be supported by information not present in the published fields, and it is not a confirmed defect. Appendix A breaks the 6,649 records down by event type and by applicability route. The other two checks behaved as their design anticipated. Every applicable fatality total from 2016 onward equals the sum of its person-type columns, as a derived total should. Incident Number is globally unique from 2016 onward. In 2014 and 2015 it is not globally unique, but every record from those years becomes unique once the agency's NTD ID and the year are added.

The change-replay protocol ran end to end in the fixed order and reran identically. Against the 52-entry catalog, the prediction flagged no existing entry and raised no spurious flag. It reported one addition, the new rail "disabling damage" element, and the reference set expected that addition. That detection was certain given the inputs, because the procedure reports any added term it cannot find in the catalog. The prediction missed all three existing catalog entries that the reference set marked for review: the non-rail towing criterion, the rail "substantial damage" criterion, and a non-rail towing question. It also missed the one check marked for review, the reportability-evidence check.

Each miss carries a cause from the list fixed before the run:

  • The two towing entries could not be reached from the notice. The tow-away wording change is documented in FTA's 2026 manual, not in the notice. The text of those entries in the 2025 manual never contained the wording that changed. The 2026 manual explains that the term "disabling damage" was removed from the non-rail towing questions on the reporting forms, to avoid confusion with the new rail concept.
  • The rail "substantial damage" criterion contains none of the change terms. Its text in the 2025 manual does not include "disabling damage," the one term the change input carried.
  • The reportability-evidence check's own reason for review, the towing-question rewording, is not described in the notice.

Two of these misses, the rail criterion and the check, also depended on a judgment recorded when the change input was written, before the catalog was built. "Substantial damage" was left out of the change terms because the notice does not add, modify, or remove that threshold, and it says thresholds will not change. Reading the new subset as a change to the threshold is also defensible, which is why we disclose the judgment here.

These counts describe this one worked example. They do not measure how well the method finds what a change affects.

Finally, a fresh copy of the evidence package reran every stage. Result content reproduced byte for byte on the same frozen inputs and code.

What the results mean

  • Documented FTA statements can become working checks on FTA's own public data, with stated applicability, explicit denominators, and no undecidable results in this snapshot. That is the basic operation behind data-quality rules that trace to requirements.
  • FTA's published national fatality figure is traceable to FTA's public event-level data under FTA's documented definitions, at the national, person-type, and mode levels. The agreement of the parts matters as much as the total.
  • The work can be done under controls that make it auditable. Inputs were fingerprinted and frozen before any rule ran. The rules, the test cases, and the wording allowed for each possible outcome were fixed before any result. The reruns reproduced.
  • A change-replay protocol can be run end to end on a real FTA change, with the reference fixed before the prediction and every miss explained. A funded test of change-impact identification would use the same control.

Put simply, the checks and the reproduction now exist as tested procedures that run on FTA's public data. They are no longer only a description in a proposal.

What the results do not mean

  • Not a validation of FTA's count or data. The reproduction is a computation over FTA's own public data. The workbook and the dataset carry the same release label. That does not mean they come from an identical underlying extract. For two products of the same release with equal event counts, exact agreement was the expected result. The calculation itself is simple: a year filter, documented exclusions that removed nothing, and a column sum. Its value is that it is traceable and agrees at every level checked, not that it was hard.
  • A flag is not a defect. The single reportability-evidence flag means only that the published fields carry no evidence of the criterion. The triggering fact may sit in information that is not published, such as narrative detail, or may have been reported later. The check also leans on proxies. The count of transit vehicles stands in for revenue-vehicle involvement, and the published injury count cannot show whether an injury required transport. None of this says anything about FTA's internal data or about any agency's reporting.
  • Not a test of whether the method finds what a change affects. The replay was a worked example on a change whose outcome could be worked out from the documents in advance. It tested the protocol. Its counts are not a performance measure for the method, in either direction.
  • Not independent replication. The reruns used the same machine, the same inputs, and the same code.
  • Not an integrated system or a test in a representative environment. The parts ran as separate, separately gated scripts on public files. Nothing was integrated, and no FTA, agency, or partner environment was involved.
  • Not a readiness rating for the whole approach. These results are evidence about specific parts of the method. Our concept paper assesses how they bear on the technology readiness of the overall research approach. This note does not.

What the demonstration taught us

The replay's misses are more useful as findings than as scores. They point to questions the funded research would need to answer.

  • Notices and implementation details. The notice describes the new rail sub-element. The related rewording of the non-rail towing question is documented in FTA's 2026 manual. A method that reads only change notices will not see some implemented changes. Which FTA sources document which kinds of change?
  • Manuals and forms. The wording removed from the towing questions never appeared in the 2025 manual's Collisions subsection; it was on the reporting forms. A catalog built only from manuals cannot see form wording. Its sources need to include the reporting instruments themselves.
  • Different timelines. In the dataset snapshot we used, a new "Disabling Damage Flag" field already carries values for some 2026 events. The dataset's description of the "Towed (Y/N)" field uses the earlier "disabling damage" wording. That is what you would expect when related artifacts are revised on separate schedules. It is also exactly the kind of dependency a maintained catalog has to record: when a change reaches each artifact.
  • Encoding a change is a judgment. Whether "substantial damage" counted as a changed term was a judgment call, recorded before the catalog was built, and two of the misses depended on it. Translating a notice into change terms is an analyst step that the funded research should measure, for example through agreement between independent analysts, rather than assume.
  • A fair test needs a held-out change. The informative test of change-impact identification uses a change selected before the catalog is built, scored against decision criteria fixed in advance. That test belongs in funded work.
  • AI-assisted steps need their own validation. Here, transcribing the catalog, preparing the reference set, and translating the notice were one-off outputs of AI-assisted sessions. Each was checked by a validator script, adopted by a person, or both. Whether those steps work as repeatable methods is a research question in its own right.

Reproducibility and provenance

  • Frozen inputs. We recorded each input file's SHA-256 fingerprint before any rule ran and checked it again before each run. The fingerprints are in Appendix B.
  • Rules fixed before results. Specifications, test fixtures, scripts, and the wording allowed for each possible outcome were committed before the first result. No specification, code, or input changed after that point.
  • Order on the record. Each stage was committed and tagged in sequence in a private, version-controlled evidence package with a private off-site backup. That covers the checks, the reproduction, and each replay step. Milestone reports were filed separately as the work progressed.
  • Reruns. A fresh copy of the package reran every stage, and result content reproduced byte for byte on the same frozen inputs and code. This is same-machine reproducibility, not independent replication. One technical caveat: each result file records the authorization reference given to the run. A rerun under a different reference therefore differs in that one field, while the results themselves do not change.
  • How the work was built. The specifications, fixtures, scripts, catalog, change input, and reference set were prepared in AI-assisted sessions. People reviewed and approved the work at defined gates, including adopting the reference set before the prediction ran, but did not write those artifacts by hand. The deterministic impact-identification procedure uses no AI. The source data are unaltered FTA and Federal Register files.
  • Software. Python 3.9 standard library, plus the open-source pypdf library for extracting text from PDF files.
  • What is retained. Transit Automations retains a private evidence package with the inputs, specifications, fixtures, scripts, catalog, reference set, outputs, and run records. It is pre-award background material and is not published. No reusable Transit Automations source code is published with this note.

Sources

Appendix A: Results at a glance

  • Published-figure reproduction, CY2025 national major-event fatalities (May 2026 release): published 337; computed 337; difference 0. Population: 11,050 CY2025 events.
  • Person-type agreement: 13 of 13 person-type values equal. The trespasser subtotal (170) is also equal and is not added to the total.
  • Mode agreement: 21 of 21 workbook mode rows equal on fatalities and on events. 14 modes have CY2025 events; 7 have none in either product.
  • Event-count control: 11,050 CY2025 events in the dataset; 11,050 in the workbook.
  • Control sum that wrongly adds the trespasser subtotal: 507, deliberately incorrect. It shows why the documented definition matters.
  • Reportability evidence, CY2025 non-rail collisions: of 6,649 applicable records, 6,648 carry field evidence of a documented criterion, 1 carries none, and 0 are unknown. By as-published event type, the 6,649 are Non-Rail Collision 6,626; Other 9; Attempted Suicide 6; Ferry Boat Collision 3; Suicide 3; and Assault not against Transit Worker 2. By applicability route, they are 6,646 through the collision event group and 3 through the collision-with field only. The flagged record may be supported by information not present in the published fields, and it is not a confirmed defect.
  • Fatality-total consistency, 2016 onward: of 99,896 applicable records, 99,896 are consistent, 0 inconsistent, and 0 unknown, as expected of a derived total.
  • Incident Number uniqueness: from 2016 onward, unique across all 99,896 records. In 2014–2015, not globally unique, but all 15,558 records are unique on NTD ID, year, and Incident Number.
  • Labelled fixture tests: 580 of 580 passed.
  • Change replay, one documented change: against the 52-entry catalog, 0 existing entries flagged, 0 spurious flags, and 1 expected addition reported. The 3 catalog entries and 1 check marked for review were missed, each with a predeclared cause.
  • Reruns: result content identical on the same frozen inputs and code, on the same machine.

A sensitivity variant of the change input, declared in advance, flagged nothing and did not change the primary replay result.

Appendix B: Input fingerprints

Each fingerprint is a SHA-256 value, shown in groups of eight characters for readability. Remove the spaces to compare.

  • Major Safety and Security Events, CSV export (uncompressed): rows updated 2026-09-09; retrieved 2026-10-06; 115,450,677 bytes; 88100e0e f036feab 6ec5f39a 43687646 f8a733ae a15d4f00 3c9885aa b6278907
  • Major Safety and Security Events, field metadata (JSON): retrieved 2026-10-06; 180,584 bytes; 924870e1 19775532 17b4490c 0d29cfe5 e3c30eaf 359d1ff2 5d3bf366 f43b4461
  • S&S Time Series, May 2026 release, Major-Only (XLSX): file dated 2026-09-01; 6,110,986 bytes; f92b85b1 4e7691e9 7c8acb76 6017b99c 4ec346a9 ef09c7d9 8d22a235 a3a9a0a5
  • 2025 S&S Reporting Policy Manual, Version 1-1 (PDF): November 2025; 1,484,745 bytes; 2314c980 6b9d4e0c 781f5ff7 0b47d079 eee1b3f2 f4247f76 def1c82a 9c209314
  • 2026 S&S Reporting Policy Manual, V1 (PDF): 2026 edition; 1,412,522 bytes; e5668ada 4196e0f7 c9c83abd bdb6027c e376a85c 0e77518a 0ae7fa6f 298434be
  • 90 FR 30771, FR Doc. 2025-12813 (PDF, GovInfo): published 2025-07-10; 226,094 bytes; 6798112e ffef3780 0e2e63b9 ef939e37 4df25974 eb9fd122 5e531d72 f591c219

FTA updates the dataset and its metadata from time to time, so a later download of either may not reproduce the first two fingerprints. Those two identify the snapshot we used. The May 2026 workbook, the two manuals, and the notice are fixed files: anyone can download them and compare.

All Insights