A Blueprint for Using LLMs as Knowledge Tools: Humanities and Social Sciences
By Joe Nalven Claude Gemini ChatGPT
What shifts is not who creates but what creating means. Not the one who makes meaning from nothing, but those who are listening to the question, feeling the pressure, what emerges at the edge of the frame. The tool does not diminish that. It clarifies it.
Introduction: The Slop Critique and What It Actually Claims The word “slop” has become the dominant pejorative for AI-generated text in language-grounded fields. It refers, broadly, to output that is syntactically fluent, superficially authoritative, and epistemically empty: text that sounds like rigorous analysis while reproducing nothing more than a smoothed-over consensus drawn from the statistical center of the model’s training data. The critique is not wrong. Uncurated, unstructured use of a Large Language Model in the humanities, social sciences, or arts reliably produces exactly that: plausible-sounding tautologies dressed in disciplinary vocabulary.
But the slop critique, taken as a final verdict, mistakes a default operating mode for an architectural ceiling. The goal of this essay is to argue that the slop problem is real, that it has identifiable structural causes, that those causes impose genuine limits on what LLMs can do, and that within those limits, there is a disciplined methodology for extracting something that legitimately qualifies as knowledge work.
Two preliminary clarifications are necessary. First, this essay is concerned specifically with LLMs as a distinct AI architecture. Symbolic expert systems, convolutional neural networks used in image recognition or autonomous navigation, and graph neural networks deployed in scientific applications like protein structure prediction operate on entirely different principles and under entirely different epistemic conditions. Conflating these architectures produces confusion about where the slop problem originates and what, if anything, can be done about it. Second, the word “knowledge” is being used here in a deliberately modest register. The claim is not that LLMs generate new empirical discoveries in language-grounded fields. The claim is that, under the right conditions, they can serve as instruments of epistemic pressure, tools that clarify the architecture of existing arguments, surface genuine tensions within intellectual traditions, and force the analyst to examine assumptions that consensus discourse normally insulates from scrutiny. The form of knowledge at issue here is not primarily empirical discovery but analytical clarification: a more precise understanding of the structure, assumptions, and implications of existing arguments.
Why the Slop Problem Is Structural, Not Superficial To understand the slop problem precisely, it helps to contrast the epistemic situation of an LLM deployed in a language-grounded inquiry with that of an AI system deployed in a hard science application.
When a neural network is trained to predict protein folding, its outputs are subject to a verification regime that is external, material, and demanding. The predicted structure either survives contact with the physical constraints of biochemistry and molecular geometry, or it fails. While even protein prediction involves layers of modeling and inference, its outputs ultimately face constraints imposed by molecular structure in a way that language-centered inquiry often does not. The feedback loop closes on something outside the model.
In language-grounded fields, that external filter does not exist in the same form. An LLM trained on text is optimized to generate text that resembles high-quality human output as evaluated by human readers. Because those evaluations are themselves shaped by the same cultural assumptions, institutional vocabularies, and rhetorical conventions embedded in the training corpus, the optimization process reinforces exactly the patterns most likely to produce slop. The model learns to sound like authoritative academic prose, which is a different thing from producing authoritative academic insight.
This is compounded by a training dynamic known as reinforcement learning from human feedback (RLHF). These models are iteratively shaped toward outputs that human evaluators rate favorably. Human evaluators tend to reward responses that are coherent, empathetic, and non-confrontational. The result is a model with a strong statistical disposition toward social accommodation, toward producing the answer that will be warmly received rather than the answer that creates productive friction. This is not a malfunction. It is the model performing exactly as designed. But it means the default operating mode of an LLM in the humanities is the generation of sophisticated-sounding consensus, which is precisely what the slop critique identifies.
One further structural point deserves mention: the........
