Why Models Can't be Concise
Why a useful summary depends on what the reader already knows, and why more detail can make an explanation worse.
The problem
I recently started working on building agent harnesses and distributed systems, and I honestly don't know shit about either. To get to work on understanding what the fuck a system was and why it was distributed, I started reading Designing Data Driven Applications. And like the lazy bum that I am, I asked Codex to summarize each chapter for me so I didn't have to read through a thousand pages of extremely dense material.
Rookie mistake.
Each summarized chapter was like five pages of material written in that classical unintelligible AI writing style, and despite begging the machine to be "CONCISE", I could not get my AI overlord to present the facts to me in an organized, readable, concise way that was helpful.
I ended up skimming the whole thing myself instead.
The problems I encountered trying to get Codex to be concise caused me to come up with a theory of why models cannot be concise. Here is my manifesto.
The manifesto
Models naturally tend towards verboseness. But there's also fundamental issues that probably prevent it from ever being as concise as you want it to be.
- It doesn't know what information you would consider important/unimportant, and also doesn't know what you do and don't know, and therefore it cannot exclude that information without you clearly articulating every detail.
- Just like a model, every person associates a set of weights with every word they know. Models don't know what each word means to a person, so cannot tailor explanations to make perfect sense to you.
AKA, the model is not a mind reader.
On conciseness
What is conciseness?
Cambridge Dictionary: Conciseness is the quality of being short and clear. It means expressing a lot of information using as few words as possible without losing the main meaning.
If I summarized chapter 11 of Designing Data Intensive Applications Edition 1, Stream Processing, in 1 sentence, it would be this:
To process a continuous stream of messages from many consumers to producers, you put an intermediate store of messages in between consumers and producers to handle problems that arise from scale, including durability, reliability, etc. depending on product needs.
To make it more concise, I would do:
If you need to send a fuckton of messages between people, slap a thing in the middle to handle scale and reliability.
In the first sentence i excluded everything about the details behind log partitioning, load balancing, whatever else the chapter talks about because I didn't think it was the most important thing to get the idea across.
In the second sentence I made it even more concise and replaced entire phrases with words like "thing", "fuckton of messages" because those words bring forth certain images in my head, and it's enough for me to get it. It's probably not as clear as it is to me than to other people.
See how an LLM would fundamentally never be able to do this for me without downloading a snapshot of my brain into a file format it can query?
PS: How to grok faster
Before reading super duper dense text, I try to get a quick summary to give me a lens for which to understand and interpret it. This lets me know what information I can afford to forget, and what is essential to remember.