CMS Stars

Behavioral Health Finally Gets a Star: The New Depression Screening Measure

CMS is adding a Part C depression screening and follow-up measure starting with the 2027 measurement year and 2029 Star Ratings, filling a long-standing behavioral health gap. Here is what plans and provider groups should do now.

In the same April 2, 2026 final rule where CMS pulled 11 measures out of the Medicare Advantage Star Ratings, it quietly added one. Starting with the 2027 measurement year and the 2029 Star Ratings, a new Part C measure tracks whether members are screened for clinical depression and, when a screen comes back positive, whether they get timely follow-up. It is the first Star measure built specifically around behavioral health, and it fills a gap that has sat open in the program for years.

A single new measure does not sound like much next to 11 removals. This one matters more than its one slot suggests, because of what it asks of you and where the work has to happen. Most Star measures can be influenced from the plan side, with outreach, reminders, and gap-closure campaigns. This one cannot. The result is produced in a clinical encounter, by a clinician, and captured as structured data, or it does not count. That changes who owns it and how early you have to start.

What CMS is adding

The new measure looks at two things. First, the share of eligible members screened for clinical depression using a standardized, validated instrument. Second, for the members who screen positive, whether they received appropriate follow-up. CMS has framed the addition as aligned with the U.S. Preventive Services Task Force, which recommends depression screening for adolescents and for the general adult population, and as filling the long-standing absence of any Star measure focused on behavioral health.

The timing gives you room. The measurement year is 2027 and the measure does not affect a published rating until the 2029 Star Ratings. CMS also framed the broader measure cleanup as a way to stop spreading attention across measures where everyone already scores well and to focus it on areas with real room to improve. Behavioral health is exactly that kind of area. The screening rates across the industry leave plenty of headroom, which means this is a measure where an early, serious effort can build a lead that is hard to copy in a single year.

Why behavioral health was the obvious gap

For all the measures the Star Ratings program tracked, it had never included one aimed squarely at behavioral health, even though depression is among the most common and most under-detected conditions in the Medicare population. The USPSTF rationale is straightforward: screening is worth doing when it is paired with systems that ensure accurate diagnosis, effective treatment, and follow-up. Screening with no follow-up is theater. That is why the measure is built as a pair, screening plus follow-up, rather than a simple screening count. It rewards the full loop: find the people who need help, then actually do something about it, on time.

For a Medicare Advantage plan, that design is a signal. CMS is not asking whether you checked a box. It is asking whether your system closes the loop on a member who screened positive. A plan can run a high screening rate and still fail the measure if the follow-up half collapses, which is exactly the kind of gap that good intentions and bad data conspire to create.

The clinical concept is not new

If “depression screening and follow-up” sounds familiar, that is because the concept is well established in quality measurement, even though it is new to the Star Ratings.

NCQA has carried a measure called Depression Screening and Follow-Up for Adolescents and Adults, known as DSF-E, in its HEDIS set for years. It looks at members 12 and older who were screened for clinical depression using a standardized instrument and, for those who screened positive, whether they received follow-up care within 30 days. The measure was adapted from earlier provider-level work and ties back to the same USPSTF evidence base. For measurement year 2026, NCQA reports DSF-E only through the Electronic Clinical Data Systems method, the structured, digital reporting path, and it now accepts instruments such as the PHQ-9 and the PROMIS Emotional Distress measure.

The new CMS Part C Star measure brings that same idea, standardized screening plus timely follow-up, into the Medicare Advantage Star Ratings for the first time. They are not the identical specification, and you should not treat them as interchangeable line for line. But the shape is the same, and that is useful, because the operational and data lessons the industry has already learned from the HEDIS depression measures apply directly to the work you are about to do for Stars.

Why this measure is genuinely hard

The difficulty is not clinical. Clinicians know how to screen for depression. The difficulty is in the data, and it shows up in three places.

First, the screen has to be captured as discrete, structured data, not as a sentence in a progress note. A PHQ-9 score of 14 written into free text is invisible to a digital measure. The same score recorded as a coded result, with the instrument identified and the value attached, counts. If your clinical documentation captures the screening as prose, you are doing the work and getting no credit for it.

Second, the follow-up has to be linked to the positive screen. The numerator is not “a behavioral health visit happened.” It is “this member screened positive, and then this follow-up occurred within the window.” That linkage, connecting a specific screening result to a specific follow-up action for a specific member, is a data problem, and it is the part that breaks when screening lives in one system and follow-up lives in another.

Third, the timing window is unforgiving. Follow-up after the window does not count, no matter how good the care was. So the measure rewards not just doing the right thing but doing it on time and recording both halves in a way the system can read.

None of these are reasons to dread the measure. They are reasons to build for it deliberately rather than assume your current documentation will pick it up. It will not.

This measure lands on two desks at once

Here is the organizational trap. The depression screening measure sits at the seam between two leaders, and seams are where accountability gets dropped.

The VP of Quality owns the rating exposure. When the 2029 Star Ratings publish, this measure is theirs to answer for. But they do not control whether a clinician screens a patient or documents it as structured data. The Chief Medical Officer and the care teams own that. They control the workflow that produces the screen and the follow-up, and they control whether it lands in the record as something a measure can count.

If quality and clinical leadership are not coordinated well before the 2027 measurement year opens, this becomes a finger-pointing exercise after the year closes, when nothing can be fixed. If they are coordinated, it becomes a rare win: a measure where almost everyone starts from a low base, the clinical work genuinely helps patients, and the plans that move first build a durable advantage. The measure is, in that sense, a test of whether your quality function and your clinical function actually operate as one system.

The practical fix is a shared definition of done. The VP of Quality and the Chief Medical Officer should agree, in writing and before the 2027 measurement year opens, on exactly what a counted screen and a counted follow-up look like in the record: which instruments, which codes, which systems, which window. Then they should watch the same dashboard, built on the same data, so neither is surprised by the other’s numbers at the end of the year. A measure that crosses an organizational seam needs a shared source of truth, or the seam becomes a leak.

What good looks like, and what silently fails

Picture two members, same plan, same year.

The first sees her primary care physician for an annual visit. The clinician administers a PHQ-9. The score is recorded as a coded result, not a sentence, with the instrument and the value attached. The score is elevated. Within the follow-up window, she has a behavioral health visit, and that visit is linked in the data to the positive screen. She counts, cleanly, in both halves of the measure. The care was good and the record proves it.

The second member gets the same care. He is screened, he screens positive, and he sees a counselor two weeks later. But his screen was typed into a progress note as free text, and the follow-up visit lived in a separate behavioral health system that never connected back to the screening result. Clinically, nothing went wrong. In the measure, he is invisible on the screening half and unlinked on the follow-up half. The plan did everything right and earns nothing for it.

The gap between those two members is not care. It is data capture and linkage. Multiply it across a population and it becomes the difference between a strong rate and a weak one on identical clinical performance. This is why the work starts with how the screen and the follow-up are recorded, long before anyone looks at a rate.

What to do now

Use the runway. Three steps over the next several quarters.

First, fix the capture before you chase the rate. Confirm that your standardized screening result lands in the record as discrete, coded data, with the instrument and score identified, every time. If it does not, no campaign will save you, because the credit was lost at the point of documentation.

Second, build the follow-up linkage. Map how a positive screen triggers and connects to follow-up across your systems, and make sure that connection survives the handoff between the screening encounter and the follow-up encounter. This is where most of the lost numerator hides.

Third, run it in parallel before it counts. Treat the 2027 and 2028 period as a rehearsal. Compute the measure on your own data, find the members who should have counted but did not, and trace each one back to the reason. Some will be genuine care gaps. Many will be capture and linkage gaps you can fix quietly, while they are still harmless.

If you already report the HEDIS depression measures, you have a head start. The capture and linkage discipline you built there transfers almost directly. If you do not, the work you do now serves both the existing HEDIS measures and the new Star measure, which is a rare case of one investment paying two ways.

And bring your network in early. If your screening and follow-up depend on contracted provider groups, the structured-capture and linkage requirements have to reach their workflows too, not just your own. A clear provider tip sheet, the right codes, and a simple feedback loop on who is and is not capturing the data cleanly will do more for this measure in 2027 than any plan-side campaign can do in 2029.

The honest version

This is the kind of measure that exposes whether your data tells the truth. A screen that happened but was documented as prose, a follow-up that occurred but was never linked to its screen, a member who got good care but shows as a gap: each one is a number that is wrong, and each one is wrong in a way you cannot explain after the year closes unless you can trace the result back to the encounter that produced it. That traceability, from the reported rate to the clinical event and back, is not a nice-to-have on a behavioral health measure. It is the difference between a rate you can defend and a rate you can only hope is right.

There is also a human stake here that the measure math can obscure. Behind every unlinked follow-up is a person who screened positive for depression and may or may not have gotten the care they needed. Getting the data right is not only how you protect the rating. It is how you know, honestly, whether you are reaching the members this measure exists to find. We would rather you be able to defend it, and to know.


HEDIS is a registered trademark of the National Committee for Quality Assurance.

See where your contracts sit against projected cut points.

Apply to our Pilot Cohort

Back to Insights