case study no.4

Assessing a data catalog’s tagging system for effectiveness

The client’s data catalog had tags, but users couldn't rely on them — an ambiguous mix of system-generated and employee-created labels was undermining search and sort. As Alloy’s sole researcher, I quantified the extent of that ambiguity and translated it into a clear path forward, which directly shaped the taxonomy rebuild that followed.

Role

UX Researcher

Timeline

4 weeks

Industry

Accounting

Team

1 researcher, 1 designer, 1 PM

Impact

Built stakeholder trust
Avoided costly misdirection
Won buy-in for user-facing work
Jump to learn more
the short version

Summary

Challenge

Search and filtering broke down because labels came from two uncoordinated sources: system-generated tags and ad hoc employee entries, with no shared logic between them. The Catalog team needed to know how big the problem actually was, and what a usable label structure would look like, before committing to a rebuild.

Opportunity

Fixing the labels also meant studying how users search in practice, specifically the mental models and shortcuts they already rely on. Those patterns could inform label structure well beyond the Catalog team's list, shaping search and discovery across the platform.

"[We] don't have very useful tags — given what we do in our space, labels don't fully support what we do."

— internal user,
Platform Catalog

how I got there

Research

RESEARCH GOALS

Evaluate Label Usefulness: Assess the effectiveness of the existing labels in helping users search and find content.

Uncover User Mental Models: Identify how users organize and apply labels in-context, and where gaps in current label list exist.

APPROACH

Methodology

n=9 participants 5 expert (Sr. Associates, Managers) 4 novice (Associate level) 3 LoS — Tax, Assurance, Advisory
1

Open Card Sort

Users create their own categories, then sort cards into them.

Research Questions

Avg. # of categories created · how sorting shifts by experience/LoS · which labels were hardest to place, and why

2

Closed Card Sort

Users sort a fixed set of new cards into Stage 1's categories, then write in and place any missing labels.

Research Questions

Avg. # of new labels written in · which categories absorbed them · how the write-ins reshape the overall structure

3

Live Label Test

After Stage 2, in a mid-fi prototype, users assign labels from the bank to sample datasets.

Research Questions

Which labels get chosen · avg. labels applied · requested-but-missing labels · reasoning behind label choices

LEARNINGS

Key Insights

Insight 01 · Golden mean

3–5 labels: the findability sweet spot

Fewer, and users couldn't find what they needed. More, and the system read as noisy instead of precise.

Insight 02 · Granularity

Specific beats broad

Experts and novices agreed: labels tied to one LoS and domain built more trust than broad, cross-LoS ones.

DELIVERABLES

Label organization trend analysis: sorted by most & least likely to be utilized based on LoS; common category groupings & trends

Refined label list: most common label requests; new labels in service areas with largest gaps; etc.

Recommendation log: solution recommendations based off heuristic score, identified gaps, and synthesized learnings

wrapping up 

Conclusion

Leadership didn’t just accept the findings - they acted on them! They assigned a dedicated librarian to expand the label list and greenlighted an internal audit of their ML capabilities to build on the research further.

Partnership deepened

Product and data teams now aligned on fixing search.

Costs avoided

Steered leadership away from a full search-engine rebuild.

Scope expanded

Approved for deeper research into tagging behavior pre-build.