A benchmark can establish that a practice is unusual. Whether being unusual is a problem is a separate question, and the distance between those two is where most benchmark decisions go wrong.
A median is a fact about whoever filled in the survey
Published comparison data gets read as though it descended from somewhere. It is assembled from practices that chose to participate, reported their own figures, and had somebody with enough time in a given month to finish a long form. What comes out the other end describes that group accurately and describes nothing else.
How it arrives does most of the damage. A median with percentiles around it looks like a scale, and a scale invites a position on it. Nothing in the data claims the middle of the distribution is where a practice ought to sit. An MGMA table is a census of the responses. There is nothing wrong with the number, and the trouble starts when it gets asked a question it was not built for.
The comparison happens in the definitions
Two figures have to mean the same thing before they can be compared, and in practice finance they usually do not. A full-time equivalent is a definitional choice: whether the hours counted are scheduled or worked, whether an advanced practice provider sits with the providers or with the staff, how a physician who works four days and takes call is represented. So is whether owner compensation sits inside staffing cost or below the line, which on its own can move an overhead figure further than any operational change made that year.
Most of the work of using a benchmark is therefore reconstructing the practice's own number to match somebody else's definition, and the answer gets decided in that reconstruction. Two people doing it from the same general ledger will not land in the same place, and neither of them has done it wrong.
Anyone who wants an outcome can find a percentile for it
A practice does not hold one position in the data. It holds several at once, because the same figure gets cut by specialty, by region, by practice size, by ownership, by whether the group is single or multi-specialty. Take a different cut and the practice moves across the distribution with nothing having changed in the building. Nobody is being dishonest. That is what a dataset with many cuts does when it reaches people who already hold a view.
A survey figure introduced to settle a disagreement usually settles it. That is the part worth watching, because the conversation then stops at the number instead of reaching why the practice sits where it does. Putting the cut on the table next to the figure, and saying which other cuts were looked at, is slower and keeps the disagreement available to be argued out.
Managing toward a number tends to move the number
The failure I have watched most often follows a comparison that was perfectly sound. Support staff per physician sits above the middle of the distribution, a vacancy goes unfilled to bring it down, and the ratio improves on the next report. The phones take longer to answer, authorization work slides later in the week, third-next-available stretches, and none of that appears on the line being managed. The figure that prompted the decision has no opinion about access.
There is a quieter version, where a deliberate choice reads back as a deficiency. A practice staffed above the median because it decided some years ago to keep waits short is looking at its own decision, returned without the reasoning attached. Whether that decision still holds is a fair thing to reopen. It is a different question from whether the practice is normal.
What it is good for is finding the question
Where this data has earned its place, for me, is as a prompt. A category sitting well away from the distribution is worth an afternoon of finding out why, and the return on the afternoon is the explanation rather than any movement toward the median. Often enough the answer has been that the definition was off and there was no gap at all. Modest distance from a median rarely repays the same afternoon.
In a compensation arrangement the data is doing a different job
One use sits outside all of the above. Where a physician compensation arrangement has to be supportable as fair market value, survey data is part of how that support gets assembled, and what is being asked of it is evidentiary rather than operational. The question is whether the arrangement can be explained by somebody else, later, from the documents.
So the care looks different. Which survey, which year, which cut, and why those were the appropriate ones become part of a record rather than part of a discussion. In most practices this is a conversation with counsel in it, and it sits apart from the monthly sort of benchmarking, which it resembles on the page and is not.