Author Archives: Kingsley

Is the race for AGI a scam? Are the tech bros con men or visionaries?

To find out, we formally introduce the Stinkometer

In research, it can be tough even for experts to tell fact from fiction. This is especially true for more speculative fields like AI. It can feel like every model is “State-of-the-Art” and that AGI (whatever that means) will be “6 months away” for at least the next 6 months. How do we know we can trust these people?

What is a visionary? What is a con man?

Visionaries sell a fringe idea of the future, usually with strong self-belief and charisma. Con men are visionaries who don’t believe what they’re saying. It’s hard to tell the difference because we can’t read minds and can’t see the future.

I solved this by adapting the insightful “New Political Compass” from Harper O’Connor, an American political YouTuber. Harper found a similar problem. He’s active in local politics and has to decide who he should build alliances with to tackle certain issues. Interestingly, the “left-right” axis wasn’t a helpful guide. Most people he canvassed were reasonable but didn’t follow politics closely enough to have a robust ideology. He realised that someone’s psychology can be more important than their stated ideology.

Therefore, it was more useful to ask: Am I talking to a reasonable person? Here, I expand his framework into the Belief Compass (or the “Stinkometer”). We can answer our question by answering three simpler ones:

Continue reading →

Will TurboQuant save us from the RAM apocalypse?

The LLM boom is causing a global shortage of the very same computer memory it needs to sustain itself. Reports suggest OpenAI’s Stargate project alone could consume up to 40% of global DRAM output. Frontier labs like Google DeepMind need to make their models more memory-efficient.


One such technique is TurboQuant, released by Google. TurboQuant is an example of an online “quantisation” method. LLMs represent information using large tensors of numerical values, where each number typically uses 64 or 32 bits. However, many values do not require full numerical precision, so we can “round” them using fewer bits and less memory. We can see this in the example below:

The rounded value now requires 4x less memory. Source

Some quantisation methods are applied offline before inference begins. TurboQuant is ‘online’ because it compresses the KV cache dynamically during inference.

Continue reading →