I wanted to test a top-tier LLM’s capability to generate an annotated bibliography for a new research topic.
In his 2026 monograph, Using Generative AI in Historical Practice, Yaniv Fox discusses two terms that he sees as integral to sophisticated use of AI by historians: agency and taste.1 Agency refers to the formulation and conception of a new research question. Taste refers to the evaluation of the LLM’s output for quality and reliability. Expertise is required for both agency and taste.
To demonstrate Fox’s ideas, I prompted a high tier LLM, Claude Fable 5.1, with a research question about the history of the fall of McCarthyism in the USA in the decades following the Red Scare. I was able to get an annotated bibliography using Claude’s Research button.
Claude Fable 5.1 began its response by identifying assumptions in the prompt. I had made an assumption that there was scholarly consensus on the rolling back of McCarthyism when in fact there are scholars who argue that pre-McCarthyism norms were not restored. Obviously, though, in the twenty years following its rise, McCarthyism ended. The congressional committees disbanded, and blacklisted people gained social stature. With my own background knowledge about the period, I know that Claude’s interpretive challenge to my framing should not be read too literally. This is an example of what Fox calls taste.
Response from Claude Fable 5.1.
Following that framing statement, Claude responded to the question with short summaries about the role of key people. After writing an initial bibliography, divided into three parts, Claude supplied ISBNs for six books and DOIs for three journal articles. Then Claude provided a button to click, called Research. The product, a separate report document generated by clicking the Research button contained over forty sources with confirmed ISBNs or DOIs.
Sample of the annotated bibliography generated through the Reserach function.
I practiced research agency by asking for a synthesis of scholarship on the rollback of McCarthyism. When checking the citations for accuracy and considering Claude’s challenge to my conceptual framing, I practiced taste.
I learned that high-tier LLMs such as Claude Fable 5.1 can offer researchers a sophisticated starting point for their questions. While LLM-generated summaries and interpretations might put limiting boundaries on a researcher’s perception of a new topic, LLMs can also alert researchers to potential additional interpretive directions. By requiring ISBNs, DOIs, or the best available citation information in the prompt, the LLM will only return reliable sources. Hallucinated sources are now very rare in the higher tier models.
Yaniv Fox, Using Generative AI in Historical Practice (Cambridge University Press, 2026), 16. ↩