By Sarah Bender
When researchers began asking Carnegie Mellon University librarians about AI-powered research tools, answering wasn't as simple as recommending one product over another.
Librarians could warn researchers about hallucinated information and fabricated citations. But deciding whether a particular tool belonged in a research workflow required more specific questions: What information does it search? How does it retrieve that information? Can users trace an answer back to its source? What are its strengths and limitations?
Before librarians could confidently guide researchers, they needed a systematic way to investigate the tools themselves.
That need led a team of CMU librarians to develop the AI-Powered Tool Assessment Framework, a structured approach for investigating AI-powered academic search tools. Initially created to build the librarians' own understanding, the framework now informs their work with researchers through the Libraries' AI in Research (AIR) program. And as librarians elsewhere confront many of the same questions, colleagues at other institutions are beginning to put the framework to use.
Seeking deeper understanding
For Huajin Wang, STEM librarian and one of the framework's creators, the project started with a practical challenge: AI-powered research products were arriving quickly, and librarians needed to understand them well enough to answer researchers' questions.

“Initially, we wanted to be able to educate ourselves,” Wang said. “A lot of librarians talk about these tools in general terms, pointing out issues with hallucinations and unreliable citations, but that doesn’t really help identify differences between specific research products. So when we came across a blog post by Aaron Tay, a librarian in Singapore, it inspired us to break down exactly what we needed to know about each tool to better advise our community.”
Doing so required enough technical depth to understand what mattered for research without turning librarians or researchers into AI engineers. The team's goal wasn't to determine the “best” tool, but to identify a consistent set of questions for understanding how individual products work, where they fall short, and whether they are appropriate for a particular research task.
From questions to method
Those questions became the AI-Powered Tool Assessment Framework, which examines four parts of an AI-powered research tool: how it retrieves information, how it generates a response, the quality of its output, and its overall usability and sustainability.

First, the framework asks what content a product searches, whether it has access to full text, and how it ranks results. Then, it examines what happens once information has been retrieved: Does the system use retrieval-augmented generation? How easily can claims in generated text be traced to cited sources?
The framework then asks evaluators to scrutinize the answer itself. Does the generated text accurately represent its sources? Does it overgeneralize scientific findings? Does it answer the user's question completely and concisely? Does it raise concerns related to bias, privacy, or copyright? A final component considers practical questions such as support for non-English searches and summaries, as well as whether vendors disclose information about environmental impact.
“Together, these elements provide a way to look beyond what a product promises and examine what happens between a researcher's query and the answer that appears on the screen,” Wang explained.
When CMU librarians applied the framework to Consensus, an AI-powered academic search tool currently being trialed by the Libraries, they found several strengths. Its corpus draws from open indexes and full-text articles as well as a growing collection of licensed scholarly material, and its filters allow researchers to narrow results using characteristics such as methodology and sample size. Wang's team found that Consensus could locate information quickly, summarize it relatively well and make it easy for users to trace claims back to sources.
The assessment also revealed limitations. Its summaries could be overly long, conclusions were often drawn from a small subset of materials, its retrieval mechanisms weren't always transparent, and results showed a preference for specific types of articles. Rather than producing a simple verdict on Consensus, the framework gave librarians a more useful picture of where the product performed well, where caution might be warranted, and what researchers should understand when using it.
That approach now supports the Libraries' broader AI in Research program. Librarians are using their assessments to provide context about individual products rather than offering unconditional endorsements, and the framework is moving into workshops and other educational efforts where researchers can learn how to ask some of the same questions themselves.
Common challenge, shared approach
At The Ohio State University Health Sciences Library, Research and Education Librarian and Assistant Professor of Practice Casey Mazzoli was confronting many of the same questions. Health sciences researchers were asking about AI tools, Ohio State was pursuing broader AI fluency initiatives, and Mazzoli was considering what librarians needed to understand before advising others.
When her director shared CMU's framework in a team chat, Mazzoli used it to evaluate Microsoft Copilot, Google Gemini, and Consensus.
“The framework offered a learning exercise to explore what elements I would want to evaluate in an AI tool,” she explained. “Working through it, I had a chance to think through aspects I wouldn’t have thought to test, and others I wouldn’t have known how to run tests for. On many criteria, all three tools scored similarly. But I found subtle differences between them and how they work, which will impact how I talk about them to researchers.”
One criterion was particularly relevant to her work in health sciences: overgeneralization. Findings in scientific literature are often specific to a particular patient population, sample, or research environment, and applying them more broadly requires additional evidence. An AI-generated summary can obscure those boundaries by making a limited finding appear more broadly applicable than the underlying research supports.
Working through that section influenced how Mazzoli evaluates AI-generated summaries. It also influenced how she thinks librarians should teach AI literacy.
“Rather than focusing on individual products that are likely to change, I think it’s important to examine features that appear across tools,” she explained. “I’ve found it’s more useful to teach students the difference between a chatbot and a search engine than to make a claim about which specific tool is more useful, and that came directly out of using this framework. Researchers can then make their own judgments and apply their own values as they evaluate various tools.”
A framework for what comes next
Librarians have long helped researchers examine sources, evidence, provenance and context when deciding what information to trust. As AI becomes embedded in scholarly research tools, that work increasingly requires understanding something about the systems finding, summarizing, and generating information.
The AI-Powered Tool Assessment Framework extends that expertise to a changing research environment. It doesn't tell every researcher to make the same decision about a product; it provides a structured way to understand what a tool is doing and make that decision more deliberately.
“As AI becomes increasingly embedded in research, we want researchers to have the knowledge to look critically at the tools they’re using,” Wang said. “The products will continue to change, but if researchers understand what questions to ask — how a tool finds information, how it generates an answer, and what its strengths and limitations are — they can make informed decisions about whether and how to use it.”