The Bard's Musings

Knowing Things vs. Knowing How to Know Things

The blue book exam promises low-tech academic honesty. But memorization alone doesn't prove understanding—and neither does generating a flawless response with AI.

I've been seeing a phrase making the rounds in education circles lately: bring back the blue book exam.

For anyone who hasn't encountered one, a blue book exam is about as low-tech as education gets. You sit at a desk with a little booklet of lined newsprint paper, a professor writes a prompt on the board, and you write your answer by hand. No Google. No textbooks. No ChatGPT. Just you, a ballpoint pen, and whatever happens to be rattling around inside your skull.

I understand the reactionary appeal. With generative AI making it nearly impossible to determine whether a student actually wrote the paper they submitted at midnight, there is something refreshingly unambiguous about saying: fine, sit down, pick up a pen, and show me what you know.

But I'm not convinced it solves the real problem.

The Illusion of Fluent Recall

I've known people throughout my life who could effortlessly dump six pages of text onto an exam page while understanding remarkably little about the underlying mechanics. Give someone a good memory, an expansive vocabulary, and a high degree of confidence, and they can produce an essay that looks brilliant on the surface without demonstrating a single shred of deep operational comprehension. The blue book catches the AI. It doesn't necessarily catch the person who has learned to perform understanding rather than develop it.

That leaves us with a more urgent question: what, exactly, are we trying to teach people to do?

There was a time when access to information was the primary limiting factor in human endeavor. If you wanted to build a circuit, diagnose an illness, or analyze a historical treaty, you needed to carry the facts in your head, own the right manual, or know which floor of the library held the right reference book. The person who had memorized the most was, in many practical contexts, the most capable person in the room. Then information became abundant, and search and retrieval became the bottleneck — knowing how to find what you needed became as valuable as having it stored internally. Now, with an AI layer capable of synthesizing and formatting that data in seconds sitting on top of an already-vast information ecosystem, the bottleneck has shifted again: validation and discernment are what the current moment actually requires.

Before accepting that framing entirely, though, the blue book advocates deserve a fairer hearing than the AI era usually gives them — because the strongest version of their argument isn't just about catching cheaters.

Cognitive psychology research on what's called "desirable difficulties" has established that the struggle to retrieve information from memory, the effortful, sometimes frustrating process of recalling something without external scaffolding, strengthens understanding and long-term retention more than looking it up does, even when the lookup produces the correct answer faster. The act of memorization isn't just about having the fact available. It's about building the neural pathways that allow you to recognize patterns, make connections across domains, and notice when something is wrong before you've consciously identified why. A person who has genuinely internalized a body of knowledge doesn't just know the answer — they have a feel for the shape of the problem that retrieval-based learners often lack. That's not nothing, and the three-tier framework I'm about to argue for is only convincing if it takes that seriously rather than dismissing memorization as an outdated bottleneck.

The Trap of Looking It Up

The opposite failure is just as real, and just as common.

We have all encountered the modern student's refrain: why do I need to memorize this if I can just look it up? That logic fails for a reason that the desirable-difficulties research makes precise: you must know something before you can recognize what you don't know. Without a solid baseline of internalized knowledge, you cannot formulate a precise, effective question. You cannot recognize when an answer is incomplete, misleading, or flatly wrong. You cannot distinguish a peer-reviewed source from someone confidently talking out of their depth. And increasingly — as anyone who has watched a confident, fluent AI hallucination get copy-pasted into a professional document can tell you — you need that foundational knowledge to recognize when a large language model is doing exactly the same thing.

The real problem isn't that the blue book is wrong or that open-book, open-network exams are wrong. It's that neither format, on its own, tests the thing that actually matters: whether a person can do something real with what they know.

A modern learning framework needs three things working together. The foundational concepts, syntax, and principles in a field that you must know well enough to recognize when something violates them — and that the desirable-difficulties research suggests you genuinely do need to struggle to encode, not just look up. The ability to navigate documentation, locate trustworthy sources, and find what you don't know efficiently when the baseline doesn't cover the edge case in front of you. And the ability to evaluate retrieved information, stress-test it against reality, and apply it to an unfamiliar problem rather than pattern-match it against a solved example you've seen before.

Closed-book exams verify the baseline actually exists. Open-book, open-network problems with access to documentation and AI tools test the retrieval and synthesis layers. But the most important part of the assessment isn't the answer — it's the audit of the process. What search terms or prompts did you use, and why? Why did you trust one source over another? How did you verify the output before implementing it? What happens to your solution if one underlying constraint changes? Can you explain the mechanics of this answer without reading the screen back to me? Those questions are exponentially harder to fake than either a memorized essay or a generated one.

The professional you want designing a bridge, treating a disease, configuring a complex network, or drafting public policy isn't the person who memorized the most static facts. It isn't the person who blindly pastes prompts into an AI and copies whatever comes back. It is the person who knows enough to spot the anomaly, knows how to find the missing piece, and — most importantly — knows enough to realize when the answer in front of them doesn't make any sense.

That's what we should be testing. We just haven't built the exam for it yet.