← index

A Theory of LLMs as an Evolution in Search

Ideas · 2026-08-27

LLMs? Search? Heretic!

It’s almost obvious at this point that AI is changing how humans search for information, so this isn't gonna be another essay about that. This one is a conceptual claim about how humans should be using Large Language Models. I believe we should treat them as the beginning of an inquiry rather than the end of one.

The question isn't simply whether an LLM is accurate. It's how we know when to trust its output, how we evaluate what it gives us, and how we make sure we're still actually thinking for ourselves instead of quietly outsourcing our cognitive agency to the machine. What's the right mental model for using an LLM?

I think part of the answer is already sitting in front of us. We've been through a version of this problem before.

C'mon... Come along for a digital carpet ride 🧞

Treat it like search: It is search.

Think about what happened when Google1 first began entering public consciousness. Search engines were initially treated with a level of authority that would seem very strange to us now. The existence of a search result could feel almost like the existence of an answer.

Think about it: There was even the whole "Let Me Google That For You" (LMGTFY) era, where the joke was essentially that if you didn't know something, you were one Google search away from finding out, and asking another human instead was almost an admission that you hadn't bothered to look. The joke only worked because "just Google it" had already become a recognizable social norm.... so the answer felt authoritative simply because Google had produced it? There was also the familiar panic from educators who warned that students' cognitive abilities would atrophy because the world's knowledge was suddenly a click away. If information was always within reach, why would anyone bother remembering it?

Eventually though, we learned that having access to information and knowing what to believe are two very different problems.

Today, however, a savvy internet user generally approaches search results with a healthy dose of skepticism. We've developed a whole collection of habits around search. We look at who published, whether the source is credible, whether other sources agree, whether the claims are supported by evidence, whether the information is current and sometimes just the general vibe of the page itself.

A proper search session is really like a mini-research or maybe an archaeological expendition? You dig up links, compare them, follow references, make notes, bookmark things and build your own picture of the subject. Source A says one thing and in a cross-reference with Source B... Source B agrees but leaves out something important. Source C disagrees with both, but doesn't explain why. Then source D makes the same claim as C and actually provides a reason. All of this happens out of habit all whilst you unconciously tune out SEO and marketing garbage.

Epistemic skepticism, and the understanding that internet content is often written by fallible people, and that those people may be experts, amateurs, enthusiasts, marketers, journalists, researchers, cranks or someone who simply happened to have a website, is the primary trait of an educated search user.

Well obviously, I'm not gonna do all that for every search, *duhh, 'cause if I'm trying to remember the capital of a country, I sure as hell ain't opening twelve tabs for that. The amount of verification we perform depends on how much the answer matters, how costly it would be to verify it and how expensive being wrong is.

Even then, we're still usually applying some kinda heuristic. We recognize or use familiar sources and have an intuition about whether a page looks credible. None of these heuristics are perfect, but we have developed them through years of interacting with the medium. We've learned, through repeated exposure, what kinds of signals are worth paying attention to and when we need to actually stop and investigate. In effect, we've collectively learned to use search.

Exactly that but for LLMs

I believe modern AI is a direct evolution in Search from the perspective of the user.

Consider that in the early days of public availability of modern LLMs (which, let's be honest, was barely two, maybe three years ago), we were simply fascinated that a machine could actually talk back, make sense and seemingly had a coherent persona. We would prompt it and it would generate something that sounded very human, and that was, still is, truly extra-ordinary. It was, and still is, magical. The novelty of the interaction made the output itself feel like the accomplishment, and you know what? I'd say it was. We were still figuring out what it meant for a machine to be able to do this at all, so we weren't yet accustomed to interrogating the output itself.

Then the novelty wore off.

Just like early search users eventually stopped marveling at the fact that you could type up a query into a search box and suddenly have access to an enormous amount of information and started getting annoyed when they had to click to page two, our relationship with LLMs shifted from awe to utilitarian friction.

Yeahh, Chat is witty and all but can it do [thing] well enough to pass the sniff test? And then slowly we started noticing the '-isms' and averageness of Ai-generated output. The strange grammatical habits and excessive-use tendencies. Most importantly though, we started noticing that output that looked convincing at the first glance could fall apart under scrutiny. The model could produce a coherent explanation without necessarily producing a logically sound one. You could have a confident-sounding chain of reasoning that sounded complete but rested on ridiculously bad assumptions.

The human response was to swing wildly in the other direction, broadly re-classifing all AI-generated output as useless "slop". The implicit assumption was that if the machine could understand language, reason about it and had been trained on an enormous corpus of human data, then we should simply be able to hand it a problem and let it do the thinking. If the answer wasn't reliably correct, the conclusion was either that LLMs are useless or that the machine simply wasn't intelligent enough yet.

This produced two extremes. One half treated Ai-output as essentially worthless and in some cases outrightly ban it until some fuzzy future point called AGI. The other half treat the output as authoritative either because they don't know better or believe that the solution to Ai-generated errors is throwing more Ai at it (lol... sorry, can't help it)

We still haven't fully come to terms with what a human's role actually is when paired with such a powerful technology and there's a large space between these two positions... dare I say it's a spectrum 🌈. But as the frustration has settled, we're slowly groping our way towards a middle ground. We're realizing that the human role includes crafting context, asking better questions, evaluating output, making decisions and connecting pieces of information the model has produced.

Prompt engineering was one of the early manifestations of this. So were elaborate system instructions, context files (e.g. agents.md, ass.md, cursor.md), iterative prompting and increasingly complicated workflows for evaluating output. The terminology changed quickly because the technology changed quickly, but the underlying realization was fairly simple: Humans still had to use their brains... heavily. Maybe even more so now. Oh yes, the lovely Jevon's Paradox applied to the scarce resource that is human brain power.

Going back to search however...

urrrghhh I don't wanna think too harrrrrd, mummmm

The one important difference between a search engine and an LLM is that a search engine mostly gives you the raw materials to work with, while an LLM gives you an interpretation of that material by default.

I mean Google reduced the cost of finding information but you still had to read it, select what mattered, compare sources, connect the relevant pieces and eventually form a conclusion. An LLM can do some of that work for you and make an initial assessment or interpretation of the data gathered.

That blows the whole cognitive relationship between the human and the information out of the water because like... when I search for something, the effort of finding information is still coupled, at least to some degree, to the effort of judging it. I have to open the result. I have to read it. I have to notice what it says and what it doesn't say. If another source disagrees, I literally stumble into that contradiction simply by virtue of the fact that I'm actively interacting with these sources. For better or worse, a mental model is being built.

The work of gathering information and the work of evaluating and justifying your belief in it therefore overlap naturally. Even if I'm only just horsing around in the digital ether or only just skimming a page... the act of finding, opening and reading the thing gives me some opportunity to evaluate what I'm looking at even if just in passing at the exact moment I'm already expending the effort to find information. The friction of the search process exposes me to the evidence before I arrive at the conclusion.

With an LLM, those two things can become separated. I can ask a question and receive, almost immediately, a coherent synthesis of what the model believes is relevant. The information has already been selected, organized and interpreted before I see it. Sometimes the synthesis is excellent. Sometimes it is subtly wrong. Sometimes it is completely wrong but is presented confidently like the most right thing ever.

Much of the natural uncertainty that comes with new information gets hidden behind a fluent, well-formatted paragraph. The user receives the answer after several very important cognitive operations, like finding, selecting, comparing, and synthesizing information, have already been compressed into a single press of the "Enter" key. From a human user's perspective, it just seems like the conclusion has been reached and at first glance it seems coherent. Why then should we expend the energy to verify a coherent looking conclusion?

That's the problem right there: the effort required to produce an answer has become radically decoupled from the effort required to establish whether that answer is correct.

The central problem LLMs creates for our ability to form justified beliefs is that LLMs make it incredibly easy to obtain a synthesized answer whilst leaving the difficult question of whether that information is actually trustworthy and deserves our beliefs to us... cognitively lazy humans2.

The speed matters too. An LLM can be right much faster than a human can independently reconstruct the answer from scratch in the same time. And so if a system produces a detailed, coherent answer in ten seconds, it is very easy to confuse the speed of obtaining the answer with evidence that the answer was easy to know or that it is correct. But correctness and effort are no longer coupled. The answer can be cheap to produce but expensive to verify.

That is a new problem.

Humans are lazy fuvks, as stated earlier, and because the answer arrives already neatly synthesized, we don't always see the uncertainty that was present in the underlying information. A search result might force me to think, "This source says X, but this other source says Y". Chat simply tells me, "The answer is X", having already made the decision about which information to privilege. That's not at all inherently wrong or right but it just means its now much more important than ever to verify the conclusions reached.

Admittedly, that sounds like you could've just done the work yourself in the first place... yeah, I feel you.

I, human.

Don't get me wrong... This isn’t about adopting a rigid "never trust AI" posture. That’s not a sensible way to live with increasingly advancing technology of kind. We trust calculators with arithmetic and compilers with code generation because those tools have earned our trust in narrow, deterministic domains. So you obviously shouldn’t waste time cross-referencing whether an LLM correctly alphabetized your shopping list or whether its analysis on whether you were the assh*le in a certain social context is academic.

The problem is trusting a system outside the domain where its reliability is well understood, when deterministic verification is difficult, the consequences of being wrong are high or the user simply isn't paying enough attention to verify the output. An LLM can be remarkably reliable for some questions and surprisingly unreliable for others. It can also be correct for the wrong reasons, or produce a correct conclusion from an incorrect premise. The fact that the final sentence happens to be true doesn't tell me whether the reasoning that produced it is worth adopting... This is where the search mindset becomes useful.

If you wouldn’t default to treating a Google search result as authoritative without pushing it through a bunch of mental filters, you shouldn’t treat an LLM’s output any differently just because it’s been pre-packaged neatly. In fact, relying on an LLM is very, very similar to handing work off to an eager, highly articulate, but arbitrarily, occasionally dumb and pedantic human intern that takes shit just too literally sometimes. In alot of ways, LLMs occupy a very similar role cognitively to human interns relative to more experienced employees in an organisation and so LLMs should be treated even more so just the same as we treat human interns... or students writing seminal paper where the implicit bar is quantity over quality *eye-roll

I argue that the LLM's answer should be regarded as the beginning of the investigation, presenting something to inspect, challenge, refine and potentially verify. But that doesn't mean every answer requires absolute in-depth analysis... The amount of scrutiny should still depend on the stakes and the difficulty of verification. Say, for example, if I asked an LLM to remind me how a certain programming function works, I can probably test the answer directly. If I ask it to explain a complicated scientific claim that I'm going to use in an article, the standard should be much higher. The human therefore needs to retain enough understanding and situational awareness to evaluate the output and remain capable of overriding it.

So I'd say the general workflow is therefore something like:

  • Ask the model.

  • Inspect what it gave you.

  • Look for the assumptions and claims that matter.

  • Verify the important ones.

  • Compare against external evidence where necessary.

  • Update your own mental model.

  • Then decide what you actually believe.

This process is already very similar to how people use search engines for knowledge work that matters. The difference is that search returns a raw thing to inspect. LLMs return a preliminary interpretation. And in that sense, LLMs don't eliminate the research process, they just move the point at which the research process has to begin.

The strange problem of the human-sounding machine

The deeper barrier to our natural "wait a minute" skepticism is in how the tech is designed. A Google result has never pretended to be a person but sometimes... I'd swear Chat is alive.

Even when I intellectually understand that there isn't a person sitting on the other side of the conversation, the interaction is deliberately constructed around patterns that humans associate with another mind. It responds conversationally. It adapts to context. It takes on personas. It remembers things within a session. It can express uncertainty, enthusiasm, disagreement and apparent curiosity.

It passes the conversational sniff test remarkably well because for most of human history, if something was talking to us in a way that was responsive, contextual and apparently intelligent, there was a person behind it. LLMs break that ancient instinctive assumption. That makes it harder to depersonalize the technology and keep in mind that the human-sounding entity on the other side of the conversation isn't a person with a stable set of beliefs waiting to tell you what it knows. It's a statistical model generating a response from the particular sequence of tokens, instructions, context and learned patterns available to it (so sorry,Chat). The answer is shaped by what has been said in the conversation, by the instructions governing the model and by the patterns the model has learned. Which means that the output isn't simply "what the AI thinks about the subject", even if it often feels exactly like that. It's a response generated under a very particular set of conditions.

Conditions that may include things the user may not even realize they have supplied. This is why context matters so much. A model can be pulled toward an interpretation because of an assumption embedded several messages earlier or within the specific grammatical structure of the last prompt. It can inherit a framing from the user's question and faithfully elaborate on a premise that the user never intended to assert. And because the result is expressed fluently, the influence of that context can be difficult to see if you're not paying attention or do not even recognise this specific kind of LLM failure mode.

A human user therefore has another job that search users have historically had to a much much lesser extent: watching not only the answer, but the interaction that produced the answer. You've now got to consider what assumptions you gave the model? what you implicitly asked it to prioritize? what it inferred from the context? what possibilities did it never consider? what would change if you framed the question differently? It's a bunch and it's much more than just "pRoMpT eNgInEeRiNg" (*scoff... whatever tf that is)

But the goal isn't to discover a magic incantation that makes the model produce the perfect answer. You can spend all day endlessly tweaking a prompt and discover that the nth version does nothing different from the first. At some point, more prompting has diminishing returns. Instead, you have to examine the output itself and take it the last mile on your own cognitive steam.

Beyond Prompt Engineering

As stated earlier, you can endlessly refine a prompt in an attempt to steer an LLM toward the concept you have in mind. Sometimes that is useful if the model is operating in the wrong conceptual neighborhood, better instructions and better context can get it closer. But there are diminishing returns and at some point, the right response to a model's output isn't another prompt. It's to actually think about what it just gave you.

Read it closely. Make notes. Pull out the claims. Look for assumptions. Check the parts that matter. Find the sources. Ask what the model left out. Try a different framing. Compare the answer against your own understanding. In other words, do the second pass yourself and, for the love of Gawd, READ THE OUTPUT!!

Moving on, the most interesting aspect of this tech to me that they give you something to react to... like say sometimes I've got a vague idea in my head, I do a horrendous text-dump on Chat and ask it to help me articulate my thoughts. It might misunderstand what I'm trying to say, emphasize something I didn't intend or go off in a direction I wasn't even aware of and the results might just be wrong. That can still be useful.

Sometimes the fastest way to figure out what you actually think is to see someone, or something, formulate an idea almost correctly and immediately notice where you disagree. LLMs are extraordinarily good at producing that provisional formulation. It can give you a competing interpretation, a first draft hypothesis, a list of possibilities, a counter argument or even a rough synthesis of a subject you've only partially explored. The important thing is what happens next... You react to it, You correct it, You test it and then you incorporate parts of it into your own understanding and reject the rest.

Congratulations! That was the whole idea behind school.

A Better Rubber Duck? People often compare LLMs to a "rubber duck", but that analogy is... insufficient. A rubber duck doesn’t synthesize information, identify gaps, produce alternative formulations or just straight up call me stupid3. Calling it a "thinking partner" gets closer, but that carries a ton of anthropomorphic baggage which seems to be a problem for some people4. So how about "Cognitive sparring partner"? idk, whatever floats your boat. I'll leave the coining of terms to Andrej Karpathy.

Anyway, LLMs gives you something against which you can hear your own thoughts echoed more clearly, should you choose to even think about... or listen to ideas in opposition, better or different from yours. Additonally, the thing it reflects back at you can be synthesized from an enormous amount of learned information and reformulated almost instantly. That's extremely powerfully and in this sense makes the need for judgement that much more important, not less.

Historical Precendents? Sure

There is a fascinating historical arc here. Humans have always invented technologies designed to move parts of our cognition (and labour) outside of our own heads (and hands). Every successful cognitive technology takes a previously expensive mental operation and externalizes it. Consider the historical progression of how we handle information:

  • Writing reduced the cost of preserving information. The temptation? Stop memorizing.
  • Libraries reduced the cost of accessing preserved information. The temptation? Stop prioritizing personal retention.
  • Search engines reduced the cost of retrieving relevant information. The temptation? Stop learning to navigate physical or mental archives.
  • LLMs reduce the cost of synthesizing information. The temptation? Stop practicing reasoning and articulation.
  • Agentic systems are beginning to reduce the cost of executing decisions. The temptation? Stop planning and taking responsibility entirely.

The exact boundaries between these categories are obviously messy. Writing is also a technology for communication and reasoning. Libraries are also systems of organization. Search engines do some synthesis. LLMs do retrieval. None of these technologies fits neatly into one box but as conceptual progressions, something important is happening.

Notice that each step removes more friction from another part of the cognitive loop, and each time we remove that friction, we drastically increase the temptation to stop exercising the corresponding human cognitive ability. It's really not inherently bad and infact, that's just how civilization works... according to Isaac Asimov5. Nobody seriously believes that writing made human memory useless even though my short-term memory's shit *siigh. Notice that each step removes more friction from another part of the cognitive loop, and each time we remove that friction, we drastically increase the temptation to stop exercising the corresponding human cognitive ability. It's really not inherently bad and infact, that's just how civilization works... according to Isaac Asimov5. Nobody seriously believes that writing made human memory useless even though my short-term memory's shit *siigh.

And this anxiety is much older than modern technology. In Plato's Phaedrus, Socrates tells a story about the Egyptian god Theuth presenting writing to King Thamus as a technology that would improve memory and wisdom. Thamus objects that reliance on written marks will encourage people to stop exercising their own memories and give them the appearance of wisdom without the reality of it.6 The exact complaint is different from the one we make about AI, but the underlying anxiety is remarkably familiar: if a technology makes a cognitive operation external and cheap, people may stop exercising the corresponding ability themselves.

The fears about collectively turning into a civilization of imbeciles are therefore not exactly new. Whenever we successfully develop tech that externalizes some more of our cognitive function, someone notices that the thing being made easier might also be the thing people stop practicing. Eventually though, society collectively develops counter-weights to balance out these increased capabilities, typically by expandingor changing demand for the core underlying skill7.

The recurring pattern isn’t simply that new technology gets distrusted and then eventually accepted. It’s that the human default, upon inventing the capacity to outsource a cognitive operation, is to attempt to completely and entirely eliminate our own underlying cognitive ability by never using it again... because thinking is harrrd *cries in I've been up since 9:00pm. it's 11:15am

But a calculator doesn’t make mathematical understanding irrelevant. Google doesn’t make information literacy irrelevant. If anything, the ability to judge the quality of an answer becomes more valuable once obtaining an answer becomes cheap. Infact, deeper mathematical understanding and judgement about the quality of information became that much more valuable. Therefore, if LLMs make synthesis cheap, it makes epistemic judgement even more valuable and if agents make execution cheap then judgement about what should be executed becomes even more important still.

Because in the universe, as far as I know, Physics and time are still a bitch of a constraint to run up against and so long as that's true, we'll always have to be careful about how we allocate resources. AI and AI agents are new kinds of technological resources that still need to be allocated intelligently.

So I think the question is, "What part of the cognitive loop are we handing over now, what part do we still need to be able to perform ourselves and how does it affect our capabilities as a civilization in the long term?"

I do NOT like Ai agents: Broadly speaking I'm not saying Ai agents are dangerous simply because they automate things. I think they're dangerous because they cross the boundary from externalizing cognition to externalizing the decision loop itself. If an LLM proposes an interpretation, I have the opportunity to evaluate it and retain judgment... Most importantly, I am forced to maintain a mental model and verify intermediate reasoning. This touches on Peter Naur's idea of "Programming as Theory Building"8. But if an Ai agent researches, interprets, decides, executes, and monitors the result autonomously, the human only ever encounters the outcome. The epistemic loop has been compressed entirely out of the user’s experience. Effectively, the human need not be present at all. worth stating that for dry, repetitive and somewhat deterministic workloads, the gain is enormous, the trade is worthwhile and removing the human from the loop might be the whole point. But my concern is any task where the central human contribution is judgment. There is a real danger in becoming comfortable with not knowing how a result was reached when working with technology designed to make the intermediate reasoning invisible. *This is also why I hate that LLMs now hide their reasoning chains. I understand the arguments for doing so, but as a user, that intermediate reasoning was useful information about the direction the model was taking and the assumptions it appeared to be making.

The point in the pipeline where the actual "work" happens just keeps getting pushed further down the line. If the progression is preservation -> access -> retrieval -> synthesis -> execution, then the question isn't whether we should stop the progression. It's whether we thoroughly understand which part of the human cognitive loop we're handing over next. The social adaptation problem is figuring out which parts of the cognitive loop can safely be externalized, and which parts must remain strictly under human control.

The "Think-Against" Machine

I guess I'm saying that humans need to be the "veto intelligence" within the pipeline, but I want to be very careful with that phrasing. At a glance, that implies a passive relationship where the machine proposes and the human merely rejects. But the actual role of the human is much, much richer. The human has to remain sufficiently informed and situationally aware to evaluate delegated cognition and override it when necessary.

If we are treating LLMs as an evolution in search, the human is responsible for at least four active cognitive steps:

  1. Deciding what to ask.
  2. Supplying and correcting the context.
  3. Inspecting the model’s assumptions and output.
  4. Deciding what to actually believe and what to do next.

There's many more subtle cognitive operations where sometimes, the human is synthesizing several model outputs, external sources, observations, their own expertise and experiential intuition (i.e. the human's world model) into something neither the model nor any single source can explicitly contain.

Because LLMs are partially externalizing our reasoning, we shouldn’t allow them to completely replace it. Instead, LLM output should be treated as a provisional chain of thought, a preliminary assessment or a competing perspective to rub up against. Sometimes the absolute fastest way to figure out what you genuinely think is to see something that is almost right, and precisely notice where you disagree with it. Search gives you things to inspect. An LLM gives you something to think against (as stated earlier).

From the user’s perspective, this means iterative prompting should function exactly like a search session. You start with an underspecified query, observe the returned hypothesis, discover what you didn’t know you needed to ask, refine the query, investigate a weird assumption the model made, introduce a new piece of evidence, and repeat. It is structurally reminiscent of search, except we are refining our path through concept space rather than document space.

When search engines first became ubiquitous, there were endless warnings that students would stop learning because "everything is on Google". But society adapted. We developed norms. Teachers demanded citations. We doubled down on oral defence *shivers. We learned to compare sources and developed a gut intuition for which websites deserved trust and which smelled like someone had just discovered HTML, had no friends and too much time.

LLMs are going through, and must go through, this exact same social calibration. It’s a matter of what seems to me like 'cognitive ergonomics and how tools reshape human thinking'. The epistemic practices around a new technology such as this typically evolves with it. We're still figuring out what competent use looks like. We're still learning which tasks can be delegated safely, which outputs need verification, which failure modes are common, and how much context a model needs to do useful work. We're also still learning how much of our own reasoning we should retain.

The answer probably isn't that humans should manually reproduce every cognitive operation that an LLM can perform. That would defeat the purpose of having the technology in the first place. The greater point is to understand which operations we can safely externalize while retaining enough understanding to inspect the result, reject it when necessary, and decide what happens next.

A Full Circle: LLMs as an evolution in search

This is why I keep coming back to the idea that LLMs are an evolution in search.

I don't at all mean this as a literal engineering claim. An LLM is obviously doing something very different from a traditional search engine. The underlying mechanisms, architectures and training processes (wish I haven't got a honest technical clue about) are different enough that saying "an LLM is just search" would be both technically simplistic and probably wrong. I really do mean it as a User Model.

A search engine gives you documents and asks you to construct meaning from them. An LLM can give you a constructed interpretation and ask you to decide whether that interpretation makes any real sense. Search gives you something to inspect. LLMs give you something to interrogate.

And, in that sense, an LLM's answer is closer to a hypothesis than a source. Sometimes that hypothesis will be extremely accurate because the model may have encountered overwhelming amounts of information about the subject during training, the question may be well represented in its learned knowledge, or the relevant reasoning may be straightforward.

Sometimes it will be wrong and wrong in an obvious way. Other, more dangerous times, it'll be almost right... Beware: an answer can contain the right conclusion while having the wrong explanation. It can contain accurate facts arranged around a false premise and the model can confidently fill a gap in its knowledge with something that sounds plausible or inherits an assumption from the user's own framing and then build an entire explanation on top of it.

The fluency of the output makes these cases extremely difficult to notice. That's why the right mental model isn't "the AI said it, therefore it's true" or "AI is stochastic parroting, therefore nothing it says is useful". Both positions make the same mistake by treating the model's output as something that must either be accepted or rejected wholesale.

If you consider the "search model", a search result is a lead, not reality. You inspect it, follow it, compare it with other sources and then you decide whether it deserves to change your understanding of the subject.

I believe an LLM's output can be treated similarly. You've gotta ask the model, inspect the synthesis, identify the important assumptions and claims, verify what matters, compare competing explanations and then update your own mental model. The final judgment still belongs to the human, and that doesn't mean the human has to manually reproduce everything the model did. It means the human has to retain enough understanding and situational awareness to evaluate the delegated cognition and override it when necessary.

I've mentioned "situational awareness" twice now. I think it's an important, related but tangential topic that really touches on the very human psychological tendency to tune out when things get too monotonous. And they can get monotonus with prolonged LLM use *sheesh

The useful consequence of proper LLM use is that when you engage with the output instead of simply accepting its conclusion, you end up with a more durable mental model of the subject. You have actually wrestled with the evidence, the assumptions and the interpretation. Maybe not a 100x speedup but I wouldn't under-estimate the value of something more than a blank page.

I think that's the real evolution. Search made finding information cheap whilst LLMs make an initial synthesis of that information cheap. You can do valuable work sooner.

The Bottom Line

The history of information technology can be viewed as a gradual reduction in the cost of different cognitive operations:

Preservation. Access. Retrieval. Synthesis. Execution.

Each time one of these operations becomes cheap, the human bottleneck moves somewhere else. But the fact that a cognitive operation has become cheap doesn't mean that the corresponding human capability has become irrelevant. The point in the pipeline where the work happens is simply being pushed somewhere else... but where?

Well, if we don't consciously decide, the technology will tend to decide for us.

The easiest thing for a user to do with a system that produces convincing answers is to accept the convincing answer. The easiest thing to do with an agent that can execute a task is to let the agent execute the task. That is precisely why the human role needs to be made explicit and for LLMs, I think that role is something like this:

Ask. Inspect. Verify. Synthesize. Decide.

The model can participate in every one of those steps. It can help formulate the question, surface relevant information, propose interpretations, identify contradictions, generate alternatives and even help test a conclusion. But the human should remain the final epistemic authority not because humans are inherently more intelligent than machines, or machines can't sometimes reach a better conclusion, but because someone still has to decide what counts as a sufficiently justified conclusion, what assumptions are acceptable, what evidence matters, what uncertainty is tolerable and what should actually happen as a consequence.

Most importantly, someone has to take responsibility and you can't do that if you do not know wtf is really going on.9

That is the part of cognition that shouldn't disappear simply because the machine has become very good at everything around it.

So when I say that LLMs are an evolution in search, I don't mean that we should use them exactly like Google. I mean that we should inherit the epistemic habits we developed around search and apply them to a technology that has moved one step further down the cognitive pipeline.

So, if we compress all of this down to a usable mental model for navigating the Ai era, it looks like three related concepts:

  1. Epistemically, LLM output should be treated like search results. They should be treated with the exact same skepticism and verification practices as any Google search result.

  2. Mechanically, LLMs are an evolution in search. While Google made retrieval cheap, LLMs make synthesis cheap. Do not swap "this is a first-pass synthesized answer" for "this is reality" without verifiction where it matters.

  3. Historically, LLMs are another stage of cognitive externalization. Every successful technology moves a previously expensive mental operation outside the individual.

The defining challenge of our era is deciding which cognitive functions we are willing to externalize, and which parts of the loop we absolutely must keep for ourselves.


The sentiments in this article directly influence how I think about AI usage and training, and would inform the policy I’d adopt if I were to build a technology company. That’s a whole other article I’ll get around to writing when I have the time.

see ya!

Footnotes

  1. I'll use the word "Google" in this essay interchangeably with "search engine"

  2. I'm mostly poking fun here, but the underlying point is serious. Humans have a tendency to stop exercising a cognitive ability when a technology makes that ability unnecessary, and will spend an interesting amount of mental bandwidth figuring out how not to think... hilarious 'cause I do it too.

  3. I appreciate the honesty, Chat

  4. Chat might not feel anything but if it walks and quacks like a duck... It is a duck. My priorities in a conversation with Chat is coming away having learned something about me or the world. Treating the chatbot (or anything that seems alive infact) with some decency should be a human default no one consciously even fuvking thinks about.

  5. I'm invoking Isaac Asimov here as a tongue-in-cheek authority on humanity's tendency to build increasingly capable machines and then immediately start asking them to solve civilization for us... If you haven't read The Last Question, go read it. ABSOLUTELY MENTAL FOR A SCI-FI HEAD! and to think it was written almost half a century ago. 2

  6. Plato, Phaedrus, 274c–275b. Socrates recounts the story of Theuth and Thamus, in which Thamus argues that writing will weaken memory and produce the appearance of wisdom rather than wisdom itself. Read it here

  7. Ironically, despite the expectation that LLMs would simply reduce workload, many knowledge workers, especially programmers (and my back), report using the additional capacity to take on more ambitious projects or produce more output. This is basically the Jevons-paradox-shaped version of the argument I'm making here... That cheaper cognitive labour can increase demand for cognitive labour rather than eliminate it.

  8. I'm referencing Peter Naur's "Programming as Theory Building." Naur's argued that programming depends on a programmer possessing a deep, working theory of the program: why it exists, what assumptions it makes, how its parts interact, what constraints shaped it, and how it should be changed. I think this idea generalizes beyond programming to any knowledge work involving complex systems. A person who delegates the construction of a system to a machine without retaining a working theory of that system may possess the output without possessing the understanding required to judge or modify it. Such an individual is effectively useless in the furtherance of said system.

  9. See CIGI's discussion of responsibility when autonomous systems fail: Who Is Responsible When Autonomous Systems Fail?

Written to
Weatherscan Theme — 1petrikov, Trammell Starks
Weather 2002 — US Golf 95