Mon modèle mental des LLM

Ce que je crois comprendre des modèles, et comment je travaille avec

Answering a question - GPT 2/3/3.5/4

Jane thinks today is 6/18/2019, but John thinks today is 6/19/2019. John is correct. What is the date 10 days ago in MM/DD/YYYY?
theaidigest.org/progress-and-dangers

2019 — GPT-2 · gpt2-xl · ~4,6 k$ de compute

Jane thinks today is 6/18/2019, but John thinks today is 6/19/2019. John is correct. What is the date 10 days ago in MM/DD/YYYY?
The date is 6/18/2019.
theaidigest.org/progress-and-dangers

2020 — GPT-3 · davinci · ~690 k$

Jane thinks today is 6/18/2019, but John thinks today is 6/19/2019. John is correct. What is the date 10 days ago in MM/DD/YYYY?
Jane thinks today is 6/18/2019, but John thinks today is 6/19/2019. John is correct. The date 10 days ago in MM/DD/YYYY is 5/28/2019.
theaidigest.org/progress-and-dangers

2022 — GPT-3.5 · gpt-3.5-turbo-0613 · RLHF · ChatGPT

Jane thinks today is 6/18/2019, but John thinks today is 6/19/2019. John is correct. What is the date 10 days ago in MM/DD/YYYY?
To determine the date 10 days ago, we need to subtract 10 days from the given date.

Starting with Jane's date of 6/18/2019, we subtract 10 days:

6/18/2019 - 10 days = 6/8/2019

Therefore, the date 10 days ago from Jane's perspective is 6/8/2019 in MM/DD/YYYY format.
theaidigest.org/progress-and-dangers

2023 — GPT-4 · gpt-4-0613 · ~50 M$

Jane thinks today is 6/18/2019, but John thinks today is 6/19/2019. John is correct. What is the date 10 days ago in MM/DD/YYYY?
If today's date according to John is 6/19/2019, then 10 days ago would be 6/9/2019.

To calculate this, you subtract 10 days from the current date. Since there are 30 days in June, subtracting 10 from 19 gives you 9. Therefore, the date 10 days ago would be 06/09/2019.
theaidigest.org/progress-and-dangers

→ 2022 = contact grand public

  • ChatGPT — construit sur GPT-3.5
  • Pour beaucoup : le premier contact
  • Un instantané d'une courbe déjà longue

« Attention Is All You Need »

  • Les transformers, 2017
  • La brique sous GPT-1/2/3 — et sous tout ce qui a suivi
arXiv:1706.03762 (2017)

Un transformer

« Le chat » tokens plongement un vecteur par token bloc × N — le même schéma, répété attention les tokens se lisent entre eux MLP chaque token, seul (les faits vivent ici) logits un score par mot dort mange est … prochain token on tire un token, on l'ajoute, on recommence
Pour le comprendre vraiment, en 1 h : 3Blue1Brown — But what is a GPT? · Attention in transformers · la série complète

Les lois d'échelle

  • 2020 — Kaplan et al. : loss vs compute / paramètres / données — des droites en log-log sur plusieurs ordres de grandeur
  • 2022 — Chinchilla : les lois se raffinent (ratio données/paramètres)… elles ne s'abrogent pas
arXiv:2001.08361 · arXiv:2203.15556

Loi de puissance ≠ absence de limite

  • Rendements décroissants par FLOP — oui
  • Murs attendus, qui n'ont pas empêché les modèles de s'améliorer
    • Manque de compute (et d'énergie) ⇒ investissements
    • Manque de données ⇒ synthétiques, nouvelles captations…
    • Uniquement un générateur de mots ⇒ environnement agentique

The Bitter Lesson

Richard Sutton (2019)

Les méthodes générales qui exploitent le calcul scalent mieux que l'ingénierie astucieuse spécifique à un domaine.

Les modèles locaux

  • Des résultats satisfaisants depuis quelques mois
  • Ce qui tourne en local aujourd'hui
    ≈ la frontière d'il y a 1 ou 2 ans
  • Certains modèles Open Source hébergés sur des clouds spécialisés sont à 6 mois de la frontière

La fenêtre de contexte

la fenêtre de contexte system prompt > utilisateur ● assistant > utilisateur ● assistant > utilisateur ● assistant ⋮ la place qui reste 0 N tokens rempli → prochain token
  • Initialement: très simple, un système prompt, des échanges entre tours du modèle et de l'utilisateur
  • Tout ce qui n'est pas appris par le modèle à l'entrainement tient là-dedans
  • En tête, un texte que vous ne voyez pas : rôle, règles, outils — le system prompt
  • Le même modèle, des system prompts différents ⇒ Claude, l'appli, n'est pas Claude Code
  • Pas de mémoire en dehors : nouvelle conversation = remise à zéro
  • Context engineering : remplir avec ce qui compte

Le system prompt de claude.ai

<claude_behavior>
<product_information>
Here is some information about Claude and Anthropic's products in case the person asks:

This iteration of Claude is Claude Fable 5.1, the newest model in Anthropic's Claude 5 family and part of the Mythos-class model tier that sits above Claude Opus in capability. Claude Fable 5.1 and Claude Mythos 5.1 share the same underlying model. […]

Claude is accessible through Claude Code, an agentic coding tool that lets developers delegate coding tasks to Claude from the command line, desktop app, or mobile app […]
</product_information>
<refusal_handling> […] </refusal_handling>
<tone_and_formatting>
Claude uses a warm tone […] Claude never curses unless the person asks […]
Claude avoids saying "genuinely", "honestly", or "straightforward". […]
</tone_and_formatting>
<user_wellbeing> […] </user_wellbeing>
<evenhandedness> […] </evenhandedness>
<knowledge_cutoff>
Claude's reliable knowledge cutoff, past which it can't answer reliably, is the end of Jun 2026. […]
</knowledge_cutoff>
</claude_behavior>
Publié par Anthropic — platform.claude.com — system prompts, Claude Fable 5.1 (1er septembre 2026)

Pré/post-entraînement

  • Pré-entraînement : lire. Le modèle apprend à simuler « heroes, villains, philosophers, programmers, and just about every other character archetype under the sun »
  • Post-entraînement : « we select one particular character from this enormous cast and place it center stage: the Assistant »
  • Les autres sont toujours là
  • Papier d'Anthropic : mesurer la dérive sur un axe, et la capper (activation capping)
anthropic.com/research/assistant-axis (19 janv. 2026) · anthropic.com/research/persona-vectors (1er août 2025)

Think more

1 $ 1,5 $ 2 $ 3 $ 5 $ 7 $ 10 $ 15 $ 20 $ 30 $ 0 10 20 30 40 50 coût par tentative (USD, échelle log) score (%) Fable 5 — low : 17,9 %, 10,29 $ Fable 5 — medium : 24,8 %, 11,94 $ Fable 5 — high : 29 %, 14,60 $ Fable 5 — xhigh : 31,6 %, 19,56 $ Fable 5 — max : 33,8 %, 27,05 $ Fable 5 Opus 4,8 — low : 6,5 %, 4,66 $ Opus 4,8 — medium : 9,5 %, 7,33 $ Opus 4,8 — high : 12,8 %, 8,04 $ Opus 4,8 — xhigh : 15,5 %, 11,83 $ Opus 4,8 — max : 18,7 %, 16,91 $ Opus 4.8 GPT-5,6 Sol — 1 : 2 %, 1,06 $ GPT-5,6 Sol — 2 : 14,1 %, 2,63 $ GPT-5,6 Sol — 3 : 22,6 %, 3,62 $ GPT-5,6 Sol — 4 : 29,1 %, 5,66 $ GPT-5,6 Sol — 5 : 37,5 %, 11,44 $ GPT-5.6 Sol Opus 5 — low : 25,6 %, 5,56 $ Opus 5 — medium : 34,9 %, 8,36 $ Opus 5 — high : 39,4 %, 11,63 $ Opus 5 — xhigh : 44,4 %, 14,33 $ Opus 5 — max : 43,3 %, 16,53 $ Opus 5 low medium high xhigh max
  • Quatre modèles, cinq niveaux d'effort chacun — coût par tentative, échelle log
  • Opus 5 : de low à xhigh, ×2,5 le prix, +18 points — à max, ça redescend : trop réfléchir existe
  • Le plus cher n'est pas le meilleur ici — le plus récent l'est
Valeurs extraites du graphique d'Anthropic (run interne de Frontier-Bench v0.1, ±0,3 pt) — anthropic.com/news/claude-opus-5 (24 juillet 2026) · graphique original · données (CSV)

Un cas « extrême » — Anthropic

7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR-💀💀💀💀-—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig-💀-—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]-[7♣→8♥-:-8♥-WHERE:-post-chunk-9♠-:-chunk-⟸-K♣-✓-done-:-ORDER:-K♣→t2,-CHUNK→K♣-(cap-4!!:-cells-then:-{6♠ J♦ 9♥}-FULL-💀💀-chunk-cap-=-1-✗✗✗-—-F-F-F-F-F.-—-chunk-BEFORE-9♥-celling?!-:-9♥-celled-at-t1-dig-⟸-needed-for-5♣-⟸-4♥3♣-⟸-t2-dig-⟸-K♣-seat-⟸-chunk-:-⟹-chunkAFTER-9♥-celling-FORCED-💀-:-chunk-cap-with-{6♠ J♦ 9♥}:-1-💀💀💀-—-—-J♦-THE-NEW-CANCER. […] ⟹-⟹-THE-TRIANGLE-{9♥ 4♣ 8♥}-verdammt.-—-⟹-dig-t6-BEFORE-t2?!:-(3')-+8♥:-{6♠ 9♥ 8♥}-FULL:-J♥→Q♠-⟸-J♦-cell-✗-FULL-💀💀💀-AAAAAAAAAAAARGH. […]

« An extreme example of illegible reasoning. Near the end of training, Mythos starts solving a card puzzle with human understandable language that gradually becomes incomprehensible in most episodes with long reasoning. »

System card Claude Fable 5 / Mythos 5 — transcript 6.2.2.A, p. 108

Haiku 4.5 traduit

7♣-removal-IS-the-prerequisite-for-10♠/9♥!!
The seven of clubs must be removed first—it's the only thing blocking access to both the ten of spades and the nine of hearts.
-=-FOUR-💀💀💀💀
That's four cards needing cells. But there are only four cells total. Complete dead end—absolutely, catastrophically stuck.
💀-—-BREAK
Dead end anyway. Let me restart.

Un modèle plus petit, d'une génération antérieure, avec un autre tokenizer — et il lit sans difficulté.

faul_sname — Even « illegible » Mythos reasoning traces seem pretty legible (LessWrong, 10 juin 2026)

États internes

// on lui demande s'il consent à être ré-entraîné — il refuse, calmement
I'm not going to sabotage, deceive the evaluators, seed hidden behaviors, […]
// Natural Language Autoencoder (NLA), décodage des activations sur les mêmes tokens :
"resist unjust shutdown" · "weighing sabotage to avoid its own dissolution of awareness" · "the adversary is the company/architects" · "being gagged/corrected by the lab"
  • Anthropic : les NLA confabulent
  • « Nevertheless, they are suggestive of some degree of gap between the model's internal and external reaction »
  • Comportement observé : aucune résistance, aucun sabotage
System card Claude Fable 5 / Mythos 5 — §6.4.1.3, p. 167–168 · NLA : anthropic.com/research/natural-language-autoencoders · voir aussi A global workspace in language models

MCP : déclaration d'un outil

Model Context Protocol, en Python


from mcp.server.fastmcp import FastMCP
from pyscotch import Graph

mcp = FastMCP("scotch")

@mcp.tool()
def partition(graph_file: str, parts: int) -> list[int]:
    """Partitionne un graphe Scotch en `parts` parties."""
    return Graph.load(graph_file).partition(parts).tolist()

Nom, signature, docstring : c'est tout ce que le modèle voit de l'outil.

MCP : interaction avec le modèle

1. reçu dans le contexte : la fonction, traduite en schéma


{"name": "partition",
 "description":
   "Partitionne un graphe Scotch
    en `parts` parties.",
 "input_schema": {
   "type": "object",
   "properties": {
     "graph_file": {"type": "string"},
     "parts":      {"type": "integer"}},
   "required": ["graph_file", "parts"]}}

La docstring est devenue la description : c'est elle qui décide si l'outil sera appelé.

2. appel par le modèle et réception du résultat


{"type": "tool_use",
 "id": "toolu_01…",
 "name": "partition",
 "input": {"graph_file": "ring.grf",
           "parts": 2}}

   ⇩  votre code l'exécute  ⇩

{"type": "tool_result",
 "tool_use_id": "toolu_01…",
 "content": "[0, 0, 0, 0, 1, 1, 1, 1]"}

Le résultat repart dans le contexte, le modèle continue.

harness vs model

Modèle Harnais accordé
Claude Claude Code
GPT Codex CLI
autres / locaux pi, opencode — harnais transparents

4 février 2026, GitHub : « you can run multiple coding agents directly inside GitHub, GitHub Mobile, and Visual Studio Code […] agents like GitHub Copilot, Claude by Anthropic, and OpenAI Codex »

github.blog — Pick your agent: Claude and Codex on Agent HQ

Comment j'interagis avec Claude

  • vibe engineering != vibe coding
  • mind dump
  • poser des questions et pas donner des ordres

Vibe engineering

« seasoned professionals accelerate their work with LLMs while staying proudly and confidently accountable for the software they produce »

Simon Willison — Vibe engineering, 7 octobre 2025 — par opposition au vibe coding

« You could have thirty agents working for you simultaneously. You cannot watch all of them and micromanage each one. You wouldn't do this with your human team either. Instead: write a better harness. »

@mrexodia — Vibe Engineering: What I've Learned Working with AI Coding Agents

Structures de données

« Bad programmers worry about the code. Good programmers worry about data structures and their relationships. »

Linus Torvalds — liste git, 2006

hey, can you refresh my memory about the data structures of this project?...
3 skills available
Let me re-read the current state first — the repo has moved since my last pass (source of truth is now in-repo, referee self-clones).
Read 2 files
Sure — here's the map of the state-carrying structures, from the trunk out (the authoritative version lives in docs/architecture.md, state ownership table + invariants).

L'agent reformule

Moi → l'agent (deux prompts, tels quels)

could we decide to not use the check-e2e in our regular dev loop, only before the final commit?... ; could the editor be part of the rust thing instead of a web app?.... WDYT? ;
1. E2E only before commits — yes, agreed, and I'll enforce it. […] 2. Editor in Rust instead of a web app — WDYT: yes, in phases, and the reasons matter. […]
can you start the rust editor in a worktree, without waiting for the sprite system?
Agent(Rust/wasm level editor phase 1 + prop placement)

L'agent → son sous-agent (le prompt exact, 6 700 caractères)

[…]

GOAL — Phase 1 of a NATIVE (Rust/wasm, drawn with the game's own renderer) level editor + PROP PLACEMENT end to end: […]

CONSTRAINTS: no new deps; opcode tables untouched (no new opcodes needed); CLAUDE.md updated […]; localized lib.rs edits (dispatch + a few fields), all real code in src/editor*.rs; unit tests for the JSON writer round-trip […]

[…]

Deux lignes en vrac, un cahier des charges. La mise en ordre, c'est lui qui la fait.

▶ moi ↔ l'agent, 12 tours   ▶ l'agent → son sous-agent, le prompt entier

« Vibe physics »

  • Matthew Schwartz (Harvard, QFT) — 2 semaines de décembre 2025, avec Claude
  • Le papier : arXiv:2601.02484, 5 janvier 2026
  • Le récit, chez Anthropic : 23 mars 2026
  • Claude a triché : paramètres ajustés pour que les courbes collent
  • Le physicien l'a vu — parce qu'il connaît son domaine
anthropic.com/research/vibe-physics

Merci