Raphael Friedmann

Understanding & using AI27 posts

A course, not an archive: numbered in reading order, start at 01. Every post stands alone if you already know the piece before it.

From zeroFoundations, 19 posts

What an LLM is, what a GPU does, and how to run one. No background needed.

01

What is an LLM, and how does it actually make words?

A jargon-free explanation of large language models: what they are, why they're basically a very good autocomplete, and how they write one word at a time.Start here
Beginner19 Mar 2026, 5 min
02

Why run AI on your own machine?

Cloud chatbots are easy and, honestly, hard to beat. The real and narrower case for running a model yourself, and the big things you give up to do it.
Beginner27 Mar 2026, 6 min
03

What a GPU is, and why AI needs one

Why running AI means buying a graphics card, what makes a GPU different from a CPU, and why the amount of memory on the card is the number that really matters.
Beginner3 Apr 2026, 2 min
04

How a 30-billion-parameter model fits on one card

Quantisation, explained for normal people: how shrinking each number in a model lets a giant fit on a desktop graphics card, and what it costs.
Beginner10 Apr 2026, 2 min
05

How to read a model's name and specs

Model names like Qwen3-30B-A3B-Q4_K_M look like a cat walked across the keyboard. Here's how to decode them, and the handful of specs that actually matter.
Beginner17 Apr 2026, 2 min
06

Mixture of Experts: how a 30B model runs like a 3B one

The trick behind Qwen3-30B-A3B: split the model into many experts, run only a few per token. Why it suits one GPU so well, and the headaches it brings.
Advanced12 May 2026, 3 min
07

How a model learns: training, in plain words

Every LLM starts as random noise, shaped by one loop: guess, measure the error, nudge billions of dials. How training works, and why it costs a fortune.
Beginner13 Jun 2026, 4 min
08

Pretraining, fine-tuning, and RLHF

A raw trained model can continue text but won't answer you. The three stages that turn it into a helpful assistant, and which one you actually need.
Advanced14 Jun 2026, 4 min
09

How a small model learns from a big one

Many of the small open models you can run at home were taught by a bigger model, not just by the internet. What distillation is, why it works so well, and what it can't hand down.
Beginner10 Oct 2026, 7 min
10

AI that isn't an LLM

Language models get all the attention, but they're one corner of AI. A tour of the other big families, what each is for, and when an LLM is the wrong tool.
Beginner15 Jun 2026, 4 min
11

What an AI agent actually is

Strip the buzzword and an agent is one simple thing: a language model put in a loop and handed tools, so it can do things instead of just talking about them.
Beginner16 Jun 2026, 4 min
12

How an agent uses tools and memory

The loop, one level down: how the model requests a tool, why its whole memory is just the growing transcript, and why long multi-step tasks fall apart.
Advanced17 Jun 2026, 5 min
13

Why an LLM trips over the r's in 'strawberry'

Ask a top model to count the letters in a word and it often gets it wrong. Not a bug, not stupidity: the model never sees letters at all. A tour of tokens.
Beginner18 Jun 2026, 5 min
14

Why you get a different answer every time

Ask a model the same thing twice and you can get two different replies. That's a deliberate dice-roll, not a glitch; one knob sets how loaded the dice are.
Beginner19 Jun 2026, 5 min
15

Why models make things up

A model hands you a wrong fact, a fake citation or an invented function with total confidence. Where 'hallucinations' come from, and how to work around them.
Beginner20 Jun 2026, 6 min
16

What the model remembers: the context window

An LLM has no memory between messages: it re-reads your whole conversation every turn, and only so much fits. Meet the context window.
Beginner22 Jun 2026, 5 min
17

How a model sees a picture

You can hand a modern AI a photo and ask about it. But a language model only understands tokens, so how is vision bolted onto a model that only knew words?
Beginner23 Jun 2026, 4 min
18

Why a model thinks before it answers

Modern models write pages of working-out before answering. That 'thinking' is computation bought with tokens: what it buys, what it costs, when it's wasted.
Beginner14 Sep 2026, 6 min
19

How to actually ask: prompting without the magic words

No secret incantations. Good prompting is just clear instructions to a brilliant, literal-minded assistant with no memory, straight from how the model works.
Beginner23 Jun 2026, 4 min

Using AI for realOps & infra, 8 posts

Twenty years of running infrastructure, pointed at using AI in practice. The honest version.

20

What an LLM can and can't do in your infrastructure

Twenty years of running systems, a couple deep in LLMs: where AI genuinely helps an ops team, and where it's a liability waiting to happen.
Ops8 Jun 2026, 5 min
21

LLMs in the terminal: a sysadmin's honest list

The concrete, everyday ways an LLM actually earns its place in a sysadmin's workflow, and the handful of rules that keep it from causing real damage.
Ops10 Jun 2026, 4 min
22

Self-hosting an LLM for your team: usually don't

Why most teams shouldn't run their own model: a small model's error rate burns more working time than the API fees it saves. Plus where it still wins.
Ops12 Jun 2026, 4 min
23

When not to use AI: an ops take

The unfashionable half of the conversation: when reaching for an LLM is the wrong choice, and a simple heuristic for telling those cases apart.
Ops14 Jun 2026, 3 min
24

When to let an agent loose: an ops take

An agent that can act can also act wrongly. After twenty years of running systems, here's how much rope I give one, and the guardrails that earn their keep.
Ops17 Jun 2026, 4 min
25

Pointing an LLM at your own documents (RAG, honestly)

Everyone wants 'a ChatGPT that knows our internal docs'. That's RAG, simpler than the hype, and most of the work is the boring retrieval half nobody demos.
Ops21 Jun 2026, 5 min
26

Prompt injection: the security hole in every LLM app

The moment your AI reads anything an attacker can influence, that content can hijack it. There's no clean fix, only containment: the honest ops briefing.
Ops24 Jun 2026, 5 min
27

Testing a model on your own work

Leaderboards tell you a model is good at leaderboards. Forty real cases from your own backlog, graded honestly, answer the question you actually have.
Ops15 Sep 2026, 6 min