case study no.4Assessing a data catalog’s tagging system for effectiveness
The client’s data catalog had tags, but users couldn't rely on them — an ambiguous mix of system-generated and employee-created labels was undermining search and sort. As Alloy’s sole researcher, I quantified the extent of that ambiguity and translated it into a clear path forward, which directly shaped the taxonomy rebuild that followed.
Role
UX Researcher
Timeline
4 weeks
Industry
Accounting
Team
1 researcher, 1 designer, 1 PM
Impact
the short versionSummary
Challenge
Search and filtering broke down because labels came from two uncoordinated sources: system-generated tags and ad hoc employee entries, with no shared logic between them. The Catalog team needed to know how big the problem actually was, and what a usable label structure would look like, before committing to a rebuild.
Opportunity
Fixing the labels also meant studying how users search in practice, specifically the mental models and shortcuts they already rely on. Those patterns could inform label structure well beyond the Catalog team's list, shaping search and discovery across the platform.
"[We] don't have very useful tags — given what we do in our space, labels don't fully support what we do."
— internal user,
Platform Catalog
how I got thereResearch
RESEARCH GOALS
Evaluate Label Usefulness: Assess the effectiveness of the existing labels in helping users search and find content.
Uncover User Mental Models: Identify how users organize and apply labels in-context, and where gaps in current label list exist.
APPROACH
Methodology
Open Card Sort
Users create their own categories, then sort cards into them.
Research Questions
Avg. # of categories created · how sorting shifts by experience/LoS · which labels were hardest to place, and why
Closed Card Sort
Users sort a fixed set of new cards into Stage 1's categories, then write in and place any missing labels.
Research Questions
Avg. # of new labels written in · which categories absorbed them · how the write-ins reshape the overall structure
Live Label Test
After Stage 2, in a mid-fi prototype, users assign labels from the bank to sample datasets.
Research Questions
Which labels get chosen · avg. labels applied · requested-but-missing labels · reasoning behind label choices
LEARNINGS
Key Insights
Insight 01 · Golden mean
3–5 labels: the findability sweet spot
Fewer, and users couldn't find what they needed. More, and the system read as noisy instead of precise.
Insight 02 · Granularity
Specific beats broad
Experts and novices agreed: labels tied to one LoS and domain built more trust than broad, cross-LoS ones.
DELIVERABLES
Label organization trend analysis: sorted by most & least likely to be utilized based on LoS; common category groupings & trends
Refined label list: most common label requests; new labels in service areas with largest gaps; etc.
Recommendation log: solution recommendations based off heuristic score, identified gaps, and synthesized learnings
wrapping up Conclusion
Leadership didn’t just accept the findings - they acted on them! They assigned a dedicated librarian to expand the label list and greenlighted an internal audit of their ML capabilities to build on the research further.
Partnership deepened
Product and data teams now aligned on fixing search.
Costs avoided
Steered leadership away from a full search-engine rebuild.
Scope expanded
Approved for deeper research into tagging behavior pre-build.