Skip to main content

Cutting Survey Results by Team Without Exposing Anyone

Why a minimum of five per chart is not enough once department, location, tenure and manager combine, and how collapsing the roster before sealing keeps a small team survey anonymous.

You can cut survey results by department, location, tenure and manager without exposing anyone, but only if the protection is applied to the combination of those attributes before a response is sealed, not to each chart after the fact. "Minimum five per chart" is where most tools stop. It's necessary, and on its own it isn't enough.

The question every HR lead asks

The overall number is rarely the interesting one. Engagement is 72% company-wide; what you want to know is whether it's 72% everywhere or 90% in Sales and 40% in Support. That means cutting the results by something, and the something is almost always attributes the organisation already knows about people: which department, which site, how long they've been here, who they report to.

Two ways to get those attributes. Ask the respondent, which costs questions and gets you the department people feel part of rather than the one on the org chart. Or take them from the roster HR already keeps, which is accurate and free, and which is where the anonymity problem gets sharp.

Why "minimum five" is the wrong unit

Most survey tools protect small groups the same way: a chart for a group of fewer than five people isn't shown. Call it the display floor. It's a real protection and we have it too.

Here's what it doesn't do. Suppose one person on the roster is the only one who is in Legal, based in Dublin, with more than ten years' service, reporting to Chen. A chart of Legal is fine, there are twelve of them. A chart of Dublin is fine. Tenure over ten years, fine. Chen's team, six people, fine. Every chart clears the floor.

But the response carries all four values together. Anyone who can read responses individually, which in most tools means the admin and the vendor, is looking at one identifiable person, however many charts were refused. The display floor guards what's drawn. It does nothing about what's stored.

This is what "k-anonymity" means in the literature: not that every single attribute is shared by k people, but that every combination of attributes on a record is shared by at least k records. Four authoritative attributes from HR make a much sharper fingerprint than four self-reported answers, precisely because HR picked them and they're correct.

Why it has to happen before the answer is sealed

In a tool where responses are encrypted in the respondent's browser and the server holds only ciphertext, there's a second constraint that turns out to be the important one.

A sealed response can't be reopened by anyone but a key holder, and it can't be re-sealed by the vendor at all. So there's no "we'll suppress it later". If the combination of attributes attached to a ballot would identify someone, the only moment to fix that is before the ballot is sealed. Whatever protection exists has to be applied at send time, or it doesn't exist.

That constraint is uncomfortable and clarifying. It rules out the usual answer, which is a filter in the reporting layer that the vendor promises to apply. A policy in the reporting layer is exactly the thing a subpoena, a breach, or a curious engineer walks around.

What we do instead

When a poll goes out to a roster, the roster is collapsed first, over the people actually being sent to.

Each attribute alone. Any value shared by fewer than five people on that roster is dropped from everyone who holds it. A department of three stops being a department for this send.

Then the combination. If the full set of values on someone's ballot would be shared by fewer than five people, the least useful attribute is removed from those people, in a fixed order: manager first, then tenure band, then location, then department. Recount, repeat, until every non-empty combination is shared by at least five. Manager goes first because it's both the finest-grained attribute and the least useful in a cross-tab, and those two facts happily point the same way.

A dropped value is absent, not labelled. The tempting move is a "Not reported" bucket. That bucket is itself a group, and it's exactly the group of people whose real value was too rare to show, which makes it the most identifying bucket on the page. So a response simply carries no value for that dimension and is counted in no bucket for it.

What survives is attached to each ballot and sealed inside the encrypted response, in the respondent's browser, beside their answers. The server never sees it in readable form. The response record has no timestamp and no link to a person, same as before.

What it costs, and telling you honestly

Collapsing before sealing has a visible cost: some cuts you'd like won't exist. A department of six where only two people answer gives a group of two, and the display floor then refuses it. That's the two floors composing, not a bug.

Two things make the cost legible rather than mysterious.

Right after sending, the admin is told what the collapse did. Either "attributes attached to 38 of 40 ballots, groups smaller than five left out", or, if nothing survived, that no attributes were attached because every group was too small. You learn it now, not three weeks later from an empty chart.

And every roster dimension on the results page carries a coverage count: how many of the responses in view carry a value for it at all. Without that line, a department chart summing to 32 of 40 reads as eight people who didn't answer. The truth is eight people whose department was too rare to keep, and those are very different stories.

The published version

When results are published as a signed document, the same rule produces a fixed catalog of cuts: one per department, location, tenure band and manager value, plus a few standard pairs like department by location. Manager is in no pair, because a manager already sits inside a department and a manager pair almost never clears five. Any cell under the floor is left out of the document entirely. The floors themselves are written into the signed document as numbers, so a reader can check every group's count against them without asking anyone.

What this doesn't do

It doesn't stop an admin who already knows the roster from reasoning about it. Nothing can; they own the list. What it stops is the product handing them a one-person cohort, which is the difference between a tool that leaks and a person who guesses.

It also doesn't authenticate the attributes. A respondent could edit their own values before submitting and move their answer between two groups that both already clear the floor. That's the same property self-reported demographics have always had, and the gain from doing it is close to nothing.

FAQ

Isn't a minimum of five per chart enough for a small team survey? For each chart on its own, yes. Not for the record behind the charts. Once department, location, tenure and manager sit on the same response, the combination can identify one person even when every chart of any single attribute clears five. The combination is what has to be protected.

Why collapse the roster before sending instead of hiding small groups in the report? Because a response encrypted in the respondent's browser can't be reopened or re-sealed later. If the combination attached to a ballot is identifying, the only moment to fix it is before the ballot is sealed. A report-layer filter is a promise; a collapse before sealing is a property.

Why is a dropped value left blank instead of marked "not reported"? A "not reported" bucket is a group, and it's precisely the group of people whose real value was too rare to show. Labelling it would make the rarest people the most visible.

Does the vendor learn anything new from roster attributes? Only that a department sits beside an email it already held, in a table our own operations role can't read. The response record still has no timestamp and no link to a person, and attributes reach the server only inside ciphertext.

Can the manager see their own team's cut? Through a published document's catalog, if the team clears five, and through a shared aggregate snapshot for self-reported cohorts. Never through a key that would let them run any cut they liked.

Read the how-to in Roster attributes and cuts, or the wider case in Anonymous feedback in small groups.

#survey results by department anonymous#small team survey anonymity#k-anonymity employee survey#roster attributes