A prior blog conveyed the availabililty of the CMS lab pricing panel decisions - here. This blog dives into the analysis. (Put on your nerd hat and thick glasses!)
##
Part 1. What This Project Does
Each summer, CMS convenes its Medicare Advisory Panel on Clinical Diagnostic Laboratory Tests, a federal advisory committee—often simply called the CDLT Panel—to review the roughly 100 new or reconsidered laboratory codes headed for the Medicare Clinical Laboratory Fee Schedule.
For each code, the Panel recommends whether Medicare should set payment by crosswalking the test to an existing priced code or by gapfilling, which sends the test into a separate process for developing a new price. The Panel’s recommendations are advisory, but they provide an unusually transparent window into how experts think about laboratory pricing.
For 2026, CMS has just published the actual vote counts for 110 laboratory codes. That creates a surprisingly rich little dataset. Rather than simply asking which tests were crosswalked or gapfilled, this analysis treats the meeting like a major sports statistics exercise: How often was the Panel unanimous? How often did members abstain? Which technologies generated disagreement? Which kinds of tests were most likely to be gapfilled? And where did superficially similar tests receive quite different votes?
The headline findings
The Panel was much more consensual than the sheer number of codes might suggest.
44 of the 110 votes were completely unanimous, and 64 were unanimous among members casting a substantive vote, allowing for abstentions.
At least one member abstained on 26 codes.
Crosswalking remained the dominant recommendation. 78 codes received a majority for a specific crosswalk, versus 31 with a majority for gapfill.
Truly split vote: Only one code failed to produce a majority for any option, making it the meeting’s clearest statistical outlier.
But the averages hide some striking patterns.
- Certain families of related tests moved almost in lockstep, while others produced sharply different votes despite technological similarity.
- Some novel technologies generated unanimous gapfill recommendations; others found convincing analogues on the existing fee schedule.
- A few votes were genuinely divided—7-5, 7-4-1, or spread across several competing crosswalks—revealing where the Panel itself appeared to have no clear view of the appropriate Medicare pricing.
The result is a kind of “Moneyball” view of laboratory policy: not just what the Panel recommended, but how strongly it recommended it, where consensus broke down, and what characteristics appear to predict crosswalking, gapfilling, or disagreement.
###
###
Part 2. Deep Dive.
All this pricing data is unusually well suited to a “Moneyball” treatment because CMS has given us a clean voting record for 110 codes, usually with 12 panelists, plus the agenda gives us useful explanatory variables: new vs. reconsidered, PLA status, code category/subcategory, and CHIM vs. MoG. The Panel’s formal job is specifically to recommend crosswalk versus gapfill for each code. The CMS agenda spreadsheet also provides the metadata fields needed to stratify the voting analysis.
I would build the analysis in about five layers.
1. The scoreboard: what did the Panel actually do?
For every code, classify the result into one of a small number of mutually understandable buckets:
| Metric | Definition |
|---|---|
| Unanimous crosswalk | Every participating member chose the same crosswalk |
| Unanimous gapfill | Every participating member chose gapfill |
| Consensus with abstention(s) | All substantive votes identical, but ≥1 abstention |
| Strong majority | Winner ≥75% of substantive votes |
| Ordinary majority | Winner >50% but <75% |
| Highly divided | Winner ≤⅔ of substantive votes |
| Knife-edge | Winning margin ≤2 votes |
| No majority | No option received >50% |
| Abstention case | At least one member abstained |
That distinction between strict unanimity and unanimity among non-abstaining members is worth preserving. An 11-0-1 vote feels very different politically from 12-0, even though the substantive consensus is identical.
And the first-pass numbers are already interesting. My parsing of the full voting PDF finds 110 code votes. Of those:
44 were strictly unanimous — everybody cast the same substantive vote, with no abstention.
64 were unanimous among those casting a substantive vote — so another 20 effectively had consensus plus abstention(s).
26 codes had at least one abstention; there were 27 abstentions total, because one code had two.
31 codes had a majority for gapfill.
78 had a majority for a particular crosswalk.
1 code had no majority at all.
The early pages illustrate the whole range nicely: X300U went 11-0 for its crosswalk, while 0609U went 10-1 for gapfill; 8XX30 split 9-2 between two crosswalks. Later, the HPV cfDNA codes were much more divided—7-4-1 and 8-3-1 in favor of gapfill.
2. Measure how strong the vote was, rather than only who won
For each code I would calculate a few simple quantitative measures.
Winner share
Winning votes ÷ non-abstaining votes.
Thus:
12/12 = 100%
11/11 after one abstention = 100% substantive consensus
10/12 = 83%
8/12 = 67%
7/12 = 58%
Winning margin
Winner votes minus second-place votes.
This catches cases where the top-line “majority gapfill” conceals genuine disagreement. A 10-2 gapfill vote and a 7-5 gapfill vote should not live in the same analytic box.
I would probably create a simple Consensus Index:
100 × winning substantive votes / all substantive votes
Then band it:
100 = unanimous
90–99 = near-unanimous
75–89 = strong consensus
67–74 = moderate
51–66 = divided
≤50 = no majority/plurality
That gives us a single continuous variable that becomes very useful when comparing classes of tests.
3. “What makes the Panel gapfill?”
This is probably the most policy-interesting Moneyball question.
We can use the agenda variables to compare gapfill propensity across:
PLA versus conventional CPT
New versus reconsidered
MoG versus CHIM
Molecular/genomic versus chemistry/immunoassay/microbiology
Oncology versus infectious disease versus rare disease, etc.
NGS versus PCR/ddPCR versus immunoassay versus mass spec
Algorithmic/MAAA tests versus analytically simpler tests
Tests with an obvious predecessor/reference code versus genuinely novel technology
For example, several clearly novel chemistry methods were unanimous gapfills: 0604U LC-MS/MS and several mitochondrial assays received 12-0 gapfill recommendations. By contrast, many MRD codes had obvious neighboring PLA comparators and therefore generated crosswalk votes—although some related MRD codes still produced very different outcomes, including 10-2 crosswalk, 9-3 gapfill, and 9-3 crosswalk results. That is exactly the kind of within-family variation worth studying.
A particularly useful output would be:
Probability of majority gapfill by test characteristic
Something like:
| Characteristic | N | Gapfill majority | % Gapfill |
|---|---|---|---|
| All tests | 110 | 31 | 28% |
| PLA | … | … | … |
| Non-PLA | … | … | … |
| MoG | … | … | … |
| CHIM | … | … | … |
| Oncology | … | … | … |
| Infectious disease | … | … | … |
| Rare disease | … | … | … |
| NGS | … | … | … |
| Immunoassay | … | … | … |
That starts to become genuinely predictive rather than merely descriptive.
4. Identify the “interesting games”
I would separately pull out perhaps 10–15 votes deserving narrative discussion.
There are several useful species:
Closest calls. The rare-disease rapid WGS code 0657U is striking: votes were spread 4, 3, 3, 2 across three crosswalks and gapfill—apparently the only case in the entire set where no choice secured a majority.
Gapfill victories over plausible crosswalks. For example, X271U was 7 gapfill versus 4 for one crosswalk and 1 for another.
Crosswalk victories despite meaningful dissent. For example, 8XX30 was 9 versus 2 between alternative crosswalks.
Abstention clusters. There is an especially conspicuous sequence of methylation tests receiving 11 votes for 0565U and one abstention across numerous clinical indications. That looks less like random noise than a panel-member-specific issue or consistent conflict/recusal pattern. We cannot infer the reason from the voting sheet, but the pattern itself is analytically interesting.
Families that move together. Some groups vote almost mechanically as a family—for example, the Borrelia/Babesia/Bartonella immunoassay series received unanimous multiplier-based crosswalk recommendations.
That lets the article distinguish individual controversial codes from entire technology classes where the Panel has developed a pricing heuristic.
5. Add several slightly nerdier Moneyball statistics
These would make the analysis much richer without becoming silly.
I would calculate a Panel Disagreement Score, perhaps 1 – winner share, so 0 means unanimity and 0.42 means a 7/12 winner. Then rank all 110 tests from least to most controversial.
I would also calculate option dispersion—how many substantive choices actually attracted a vote. A 7-5 vote has two schools of thought; a 4-3-3-2 vote has four. Those are different kinds of disagreement.
Then add family consistency. If five technologically similar tests are presented together, how often does the panel make the same methodological decision? The DNA-methylation series and the rare-disease sequencing series would be excellent examples.
Finally, there is a potentially very useful “reference-code gravity” variable: when CMS/presenters offer an obvious established comparator, how often does the Panel take it rather than gapfill? A more advanced version could ask whether the Panel tends to accept simple 1:1 crosswalks more readily than multiplier or arithmetic crosswalks such as “code ×2” or “A + B.” The results sheet contains many such compound recommendations.
So the eventual headline dashboard might read something like:
2026 CDLT Panel: 110 Votes in Two Days
71% majority crosswalk
28% majority gapfill
40% completely unanimous
58% substantively unanimous including abstentions
24% involved at least one abstention
9% had winner support of two-thirds or less
1 vote produced no majority
Those percentages are based on my initial extraction; before publishing I would run the finished structured dataset against the PDF one more time and lock the definitions.
###
The really fun second stage is to turn the agenda and voting PDF into a 110-row spreadsheet, one code per row, with perhaps 25 variables and automatic scores. Then we can answer things like, “Are MoG codes twice as likely to get gapfilled?”, “Are PLA codes actually more contentious?”, “Which technology families have the greatest disagreement?”, and “What characteristics predict a 12-0 crosswalk?” That would indeed be Moneyball for CLFS policy nerds.
###
OK
###
Yes — I went ahead and built the 110-code Moneyball dataset. The agenda is especially valuable because Appendix 3 supplies the structured explanatory variables for every code: code type, category, subcategory, subcommittee, and full descriptor. I matched those 110 agenda records to all 110 voting records and verified that every ballot’s option counts reconcile to its reported total responses. The voting file, for example, distinguishes competing crosswalks, gapfill, and abstention vote-by-vote.
Download the 110-code CDLT Moneyball workbook
It contains five sheets: Dashboard, Code Data (one row per code, ~40 variables), Group Summary, Split Votes, and Definitions. The quantitative scores—winner share, disagreement score, unanimity, abstention flag, highly-divided flag, crosswalk/gapfill totals, etc.—are formula-driven so the workbook functions as an analytic model rather than a static table.
A few findings jump out immediately. MoG was not more likely than CHIM to gapfill: 19/75 MoG codes (25.3%) versus 9/31 CHIM codes (29.0%). But MoG was distinctly more contentious: its average winning share was 89.2% versus 97.2% for CHIM, and all 10 codes meeting our “highly divided” threshold were MoG codes.
PLA status looks interesting. There were 99 New PLA codes, and 31.3% went majority-gapfill. None of the seven ordinary “New” codes or four reconsidered codes did. New PLA codes also had lower average winner support (91.1%) than ordinary new codes (93.8%) or reconsidered codes (97.9%). That is suggestive rather than causal, given the tiny comparison groups.
The technology-family contrasts are even more fun. Chemistry had a high gapfill rate—5/9, or 55.6%—but almost no controversy: average winner support was about 98%. In other words, the Panel often said “there is no good crosswalk” very confidently. By contrast, whole-genome sequencing was the battlefield: nine WGS codes averaged only 70.4% winner support, and the sole ballot with no majority at all was a rapid-WGS code.
And there is a beautiful methylation contrast. The source-defined “Methylation Profile” group had 14 codes, zero majority gapfills, and about 97% average winner support. The six “Methylation, dPCR” codes went 6-for-6 to gapfill, with considerably more dissent. That is exactly the sort of pattern that would disappear in a simple tally of “molecular tests.”
###
This is now a good platform for Part 3: we can turn the spreadsheet results into the polished narrative analysis, with perhaps 8–12 genuinely interesting “Moneyball findings” rather than drowning the reader in 110 individual votes.
Alarmingly detailed 5-page Excel:
###
Snapshot of the just-linked excel:
| omg five worksheets of detail from AI |
OK
###
Part 3
###
Part 3. What the Votes Actually Tell Us
Once the CMS agenda is joined to the voting results, the meeting becomes more interesting than a simple count of crosswalks and gapfills. The agenda supplies a remarkably useful set of variables for each code—new versus reconsidered, PLA status, code category, subcategory, CHIM versus MoG, and the full test descriptor. The voting file then supplies the actual distribution of votes among competing crosswalks, gapfill, and abstention.
I analyzed all 110 codes using both sources. Several patterns stand out.
1. Crosswalk wins most of the time—but that understates how selective gapfill is
Of the 110 codes, 78 received a majority recommendation for a particular crosswalk, 31 received a majority for gapfill, and one produced no majority at all. Thus, roughly 71% went crosswalk and 28% gapfill.
That establishes a useful baseline. The Panel is not reflexively sending novel tests to gapfill. Most of the time it can find an existing CLFS analogue that a majority considers adequate. But almost three tests in ten were judged sufficiently different that the Panel preferred creating a new price.
2. The Panel is strikingly consensual
The most surprising top-line statistic may be how rarely these experts truly disagree.
44 of 110 votes were strictly unanimous—every participating panelist selected the same option, with no abstention. 64 were unanimous among members casting substantive votes, allowing for one or more abstentions. The mean winning vote share across all 110 codes was about 91.5%.
In other words, the typical code did not produce a lively philosophical debate between crosswalking and gapfilling. Most decisions were fairly decisive.
That makes the minority of genuinely divided votes more informative.
3. Nearly all serious disagreement occurred in MoG—not CHIM
The agenda divides work chiefly between CHIM—Chemistry, Hematology, Immunology and Microbiology—and MoG, Molecular Pathology and Genomic Sequencing.
Interestingly, the two groups were not dramatically different in their propensity to recommend gapfill:
MoG: 19 of 75 codes, 25.3%
CHIM: 9 of 31 codes, 29.0%
So MoG was actually slightly less likely to recommend gapfill.
But MoG was much more contentious. Average winner support was only 89.2% in MoG, compared with 97.2% in CHIM. More strikingly, using a relatively generous definition of a “highly divided” vote—winner receiving no more than two-thirds of substantive votes—all 10 highly divided votes were MoG cases.
That may be one of the strongest Moneyball findings. The difficult policy question is not simply “molecular tests get gapfilled.” Rather, complex molecular/genomic tests are where experts have trouble agreeing on the right analogue.
4. PLA codes account for essentially all the action
There were 99 New PLA codes in the dataset. Among them, 31—31.3%—received majority gapfill recommendations.
By contrast, none of the seven conventional “New” codes and none of the four reconsidered codes received a majority gapfill recommendation.
The small comparison groups mean this should not be treated as a formal causal finding. But practically, it makes sense. PLA codes are frequently created precisely because the laboratory service does not fit comfortably into the older generic CPT taxonomy. Those are therefore the cases in which pricing analogies are most likely to become difficult.
They also generated more disagreement: average winner support among New PLA codes was about 91.1%, compared with roughly 93.8% for ordinary new codes and 97.9% for reconsidered codes.
5. Chemistry frequently went to gapfill—but with remarkably little argument
This is a nice counterexample to the idea that gapfill means uncertainty.
Among the nine codes classified simply as Chemistry, five—56%—received majority gapfill recommendations. Yet average winner support across those codes was nearly 98%.
Several were emphatic 12-0 gapfills. The LC-MS/MS bradykinin test, for example, was sent to gapfill unanimously, as were several mitochondrial-disease assays using PAGE, Western blot, or spectrophotometric methods.
The interpretation is important:
Gapfill can represent certainty, not confusion.
Sometimes the Panel appears to be saying, quite confidently, “there simply isn't an appropriate existing CLFS analogue.”
6. Whole-genome sequencing was the real battleground
The nine codes categorized as WGS were unlike almost anything else in the meeting.
Their average winning vote share was only about 70%, versus 91.5% overall. Six of the nine met the “highly divided” threshold.
Several produced classic Moneyball-style close calls:
X271U: 7 gapfill versus 5 divided among crosswalk alternatives
X302U: 7 gapfill versus 5 crosswalk
X291U: 7 for a crosswalk versus 5 gapfill
0658U: an 8-vote winning crosswalk
0659U: another 8-vote winning crosswalk
And then came 0657U, the outlier of the entire meeting.
That rapid-WGS comparator code drew votes of 4, 3, 3, and 2 across four substantive options. No option received a majority. It is the only such case among all 110 codes.
That is probably the single best “Moneyball” anecdote in the dataset: 110 tests, one genuine hung jury—and it was rapid whole-genome sequencing.
7. MRD shows that apparently similar technologies need not price alike
The molecular residual disease group is particularly revealing because several codes look superficially related yet generated very different recommendations.
For three single-time-point solid-tumor MRD codes, the Panel strongly favored crosswalk to 0307U: two were 12-0 and another 11-1.
But the baseline MRD codes fractured:
0641U: 10-2 crosswalk
0646U: 9-3 gapfill
X278U: 9-3 crosswalk
The underlying voting record confirms those contrasting results.
Meanwhile, the leukemia MRD codes 0645U and 0644U were both 12-0 gapfill recommendations.
This is an important lesson for companies attempting to predict CMS pricing from neighboring codes. “It's another MRD test” is not enough. Specimen, baseline versus surveillance use, assay architecture, and available comparator codes can move the vote in entirely different directions.
8. The methylation codes produced two almost opposite pricing patterns
Another unusually clean natural experiment appears in the methylation tests.
A set of 14 tests classified by the agenda as Methylation Profile received zero majority gapfill recommendations. Their average winning share was about 97%, generally favoring established molecular crosswalks.
Indeed, a long sequence of disease-specific methylation tests repeatedly received 11 votes for crosswalk to 0565U, generally accompanied by one abstention.
By contrast, the six codes categorized Methylation, dPCR went six-for-six to gapfill. Several received 10-2 or 11-1 gapfill votes. The smoking/alcohol/lung-cancer series, for example, repeatedly rejected crosswalk to 0566U in favor of gapfill.
This is stronger than saying “methylation tests tend to X.” They don't.
The pattern suggests that assay architecture may matter as much as the biomarker concept itself. A broad methylation-profile test may have an accepted pricing analogue while a focused methylation-sensitive digital-PCR service may not.
9. Abstentions were concentrated rather than random
There were 26 codes with at least one abstention, representing 27 abstention votes altogether.
But abstentions were not evenly scattered. They occurred overwhelmingly in MoG: 24 of the 26 abstention-containing codes were MoG cases.
One especially conspicuous pattern is the repeated methylation-profile series, where many successive codes received 11 votes for the same crosswalk and one abstention. For example, the Parkinson disease methylation code 0626U shows exactly that 11-to-0565U plus one abstention pattern.
The public vote tally does not identify why a panelist abstained, so it would be inappropriate to infer conflict of interest or another specific reason. But analytically, repeated abstentions across a technology family are different from sporadic disagreement and should not be counted as evidence that the Panel was uncertain.
That is why distinguishing strict unanimity from substantive unanimity matters.
10. Infectious-disease panels were almost boringly easy
At the opposite extreme from WGS, the eight codes in the agenda category Microbiology; Infectious Disease panel went 8-for-8 to crosswalk, with 100% average winner support.
The Borrelia/Babesia/Bartonella immunoassay group is almost comically tidy. The Panel unanimously supported arithmetic crosswalks such as 0042U ×2, 0041U ×4, or 0041U ×5.
This also shows that the Panel is perfectly comfortable with constructed or multiplier crosswalks when it believes the existing fee schedule supplies good building blocks.
Novelty alone therefore does not force gapfill.
11. “Highly divided” is a much more useful signal than simply “not unanimous”
There were many technically non-unanimous votes that were hardly controversial—11-1, 10-2, or 11-0-1.
Using winner ≤ two-thirds of substantive voters identifies only 10 genuinely difficult cases.
Those 10 disproportionately involve:
WGS,
hereditary genomic panels,
cfDNA/ddPCR,
and closely contested gapfill-versus-crosswalk choices.
That is probably a better list for policy attention than all 66 votes that failed strict unanimity.
The distinction is analogous to baseball analytics: the useful variable is not simply whether a batter made an out; it is how much information that particular event carries.
12. The larger lesson: the Panel seems to price by “analogy confidence”
The strongest pattern across the 110 votes may not be any particular technology.
The apparent decision rule is closer to this:
When experts can identify a credible existing laboratory service whose resource profile and technical structure seem analogous, crosswalk usually wins—and frequently wins overwhelmingly. When the analogy becomes strained, gapfill becomes much more attractive. When several plausible analogies compete with one another, the vote becomes divided.
That framing explains several otherwise puzzling combinations:
Chemistry: high gapfill but high consensus
Infectious-disease panels: near-universal crosswalk and near-universal consensus
WGS: mixed outcomes and very low consensus
Methylation profiles versus methylation dPCR: same broad biological concept, opposite pricing behavior
MRD: similar clinical label, but different pricing decisions depending on assay and use case
So perhaps the most useful “Moneyball statistic” isn't simply probability of gapfill.
It is confidence in the pricing analogy.
And the voting distribution gives us a surprisingly good proxy for that confidence.