Classification Literacy: The AI Skill Library School Taught Me For AI-Assisted Qual Analysis
What an AI "theme" or a "category" leaves out turns out to be exactly the kind of question discussed in library school.
In my first semester at grad school as an Information Studies/Library Science student, a professor had our class debate what counts as a sandwich — is a panini one? A hot dog? We worked through competing definitions, including the idea that a sandwich is one thing contained within two others: a slice of ham between two pieces of bread, or, in one extreme version we landed on, “a person laying in bed, situated between their blanket and bed.” It was useful groundwork: it trained us to notice how much work a boundary does before anyone questions where to catalog something.
That instinct resurfaced a few weeks ago, while I was reviewing interview themes an AI tool had pulled from a batch of transcripts. The themes looked reasonable. They had clean labels, tidy summaries, and nothing obviously wrong. But one category bothered me: it folded two different user behaviors into a single label. I couldn’t say why until I recognized the question I was asking — the one that UBC’s School of Information had trained in me: whose/which experience did this category treat as the default, and whose/which as an exception to it?
I don’t think that’s a question I’ve seen much in AI literacy advice. It’s the kind of question that training at a library school would give you, about something older than AI: the classification system. So here’s something I think is worth discussing: what does it take to interrogate an AI’s classification decisions for AI-assisted qualitative thematic analysis?
What a Information Studies/Library Science Degree Actually Teaches You to Ask
Whenever I mention my background in information studies and library science, most people picture someone helping you find a book, or cataloging information into a database system. This is only a narrow case of a skill that the degree actually trained me with. Every library science program accredited by the American Library Association (ALA) requires coursework in classification and cataloging as a core competency, because the field treats “naming classifications” as an exercise of power, not just a neutral administrative task.
For example, Geoffrey Bowker and Susan Leigh Star’s foundational work on classification systems “argues no scheme is neutral”: every category and standard elevates one way of seeing the world while erasing others. If you break this down, a category does four things at once: it draws a boundary, picks exemplars that stand in for the “typical” member, creates a catch-all for whatever doesn’t fit, and “drifts” as it gets reused. Hope Olson’s analysis of the Library of Congress Subject Headings and the Dewey Decimal Classification shows what happens when nobody checks that “drift”: categories that have organized libraries for over a century have treated women, people of color, and other groups outside the system’s original frame as secondary or missing altogether. Naming, in her account, is never just descriptive, but rather a decision about whose experience becomes the norm, and whose becomes the residual case.
In a lot of contexts with AI product development, user trust, transparency, and intent are treated as an underspecified word until UX research decomposes it down to specifics. So in a sense, I see “category” as something that a library student or UX researchers spends time decomposing before we build one ourselves.
What That Question Catches in an AI-Generated Theme
I see a parallel between my graduate training of cataloging and classification, and thematic analysis for UX research studies. Let’s talk about it with an example: in a user test, we may see a theme called “confusion of what AI can do” from the majority of users. Underneath it were actually two groups of behaviors: participants who misunderstood the AI’s capabilities because the AI product hadn’t communicated to them clearly, and participants who deliberately tested its limits in an unorthodox (e.g., asking edge-case questions, testing different model capabilities, etc.) to calibrate trust before relying on it for something that mattered. If we use an AI LLM tool to do thematic analysis on the qualitative data, maybe both groups of participants get folded into the theme of “confusion,” because both showed an outcome of being unsure about what the AI could do. But actually, the second group wasn’t confused; they were probing the AI further as what any careful person would do before trusting a new AI tool. Perhaps in this case, the AI treated the first group’s confusing experience as the “default group”, and merged the second group into the same theme since participants in the second group showed the same confused outcome too.
I feel like this could be an example that’s worth noting, and splitting those two groups apart is a kind of “construction” needed by the UX researcher. In this hypothetical example, I would draw a different boundary than the AI did and separate the two groups. That would not be as a neutral correction of its error, but rather an injection of the sense of scrutiny and in-depth observation that I bring to cataloging the types of human behavior and user intent from my experience as a human UX researcher.
So, I think esssentially the split between the two groups isn’t arbitrary or neutral at all. From select literature on qualitative coding using AI tools, one study testing LLMs against inductive thematic analysis found they reliably surface a set’s main, explicit theme (e.g., “users don’t know what the AI can do”) but missed the latent distinctions within it. Though that’s only one study’s finding, it lines up with a living concern from information science professionals, where a 2025 conference analysis found automating classification carries real risks around accuracy, transparency, and bias; essentially arguing that human expertise becomes more necessary in these cases, not less.
I don’t think any of this means the AI’s theming was wrong though. It only means treating a theme as evidence without asking whose experience got treated as “default” is where the actual research risk lives.
Why I Think Classification Literacy Is Also Important
Most of the AI literacy advice I’ve come across for researchers is prompting technique: how to phrase a request so the model returns something more useful. That’s a real skill, but I also think that “category-formation literacy,” or the habit of asking whose experience a boundary treats as the norm before accepting it as a finding, is also essential when it comes to doing thematic analysis with AI tools.
So, I feel like the useful version of category-formation literacy wouldn’t be auditing an AI’s themes after the fact, but building the coding structure (e.g., coding decriptions, codebook, etc.) into how the team designs the coding process before the AI output gets treated as evidence. And an experienced UX researcher would be able to bring in their human judgment for this into the AI tool to make sure that the themes make sense for each indivdiual project that they work on by following the specific research requests for each individual product team.
A Step Worth Adding to AI-Assisted Coding
I don’t think you would need a library science degree for this, but it would be worthwhile to add one step to how your UXR team handles AI-assisted coding. Before any AI-generated theme gets treated as a finding, it would be worth it to have a human in the loop. So in the synthesis session, have a UX researcher interrogate the AI’s themes against three questions:
Who decided this boundary between themes, and on what basis?
What’s the catch-all category quietly absorbing (in other words, what “other” or “miscellaneous” cases are mostly hidden in the data)?
Does this scheme assume one dominant way of users experiencing the problem? Or are there other iterations of the themes that the AI can also suggest?
I feel like that’s a small addition to a coding workflow, and it doesn’t reject the AI’s output outright. It requires treating a category like any other research construct: built on a set of carefully made decisions to preserve rigor and validity.
The Sandwich Debate Never Settled Anything, and That’s the Point
So, going back to the example I shared about the sandwich debate in library school…in the end, it never actually settled anything. Nobody left that classroom agreeing on whether a hot dog or a human lying in bed counts as a sandwich. And that was the point.
The exercise wasn’t about reaching a tidy definition; it was about noticing, before you accept a boundary for a classification, whose case sits at the center of it (as the default case) and whose sits at the edge and gets folded in. That’s the same question I’d ask of an AI-generated theme.
So what does it take to interrogate an AI’s classification decisions for qualitative thematic analysis? I don’t think the answer is a sharper prompt, or a more careful read-through once the themes are already sitting in a deck. It’s treating category-formation as part of the research design itself, something a UX researcher builds into the coding process before an AI’s theme gets to stand in as a finding. An Information Studies/Library Science degree happened to train that habit into me years before I had any use for it with AI tools. But you don’t need the degree to ask the question. You just need to notice that a theme, like a subject heading, is never only a description. It’s a decision about whose experience got to be the norm, and whose got folded in as an exception to it.

Notes: This piece reflects my own reading of the cited sources. If you’ve applied a classification or taxonomy background to evaluating AI output and have a different experience or notes to share, I’d love to hear about it!
Thanks for reading Signals to Solutions! Subscribe for free to receive new posts and support my work.
Sources Referenced
Classification theory — library & information science:
Bowker, G. C., & Star, S. L. — Sorting Things Out: Classification and Its Consequences (MIT Press)
Olson, H. A. — The Power to Name: Locating the Limits of Subject Representation in Libraries
American Library Association — Core Competences of Librarianship (2023 policy)
AI, classification & evaluation:
De Paoli, S. (2024). “Performing an Inductive Thematic Analysis of Semi-Structured Interviews With a Large Language Model.” Social Science Computer Review, 42(4), 997–1019.
Cheng et al. (2025). “Reimagining Knowledge Organization with AI and Human-in-the-Loop.” Proceedings of the Association for Information Science and Technology.




留言