Artificial Intelligence

Artificial intelligence is a production tool here before it is a topic of debate, and that is what gives this category its angle: the articles come from systems actually built, not from demos. You will find the full RAG chain, from choosing an embedding model to the evaluation metrics that lie, the shift to CAG once the context window explodes, agent and MCP architecture, persistent memory and Claude Code hooks, prompt caching and what it really does to the bill. You will also find criticism of that same ecosystem when it slips: subscriptions sold as unlimited and then rationed, token consumption industrialized, AI text detectors that should never have existed, a French-language hype bubble that talks a great deal and ships very little. And at regular intervals, the questions the engineering does not settle: what automation takes away from the person writing the code, what a model makes the dead say, where the line runs between a tool and a crutch. Written by someone who deploys these systems in production and publishes the failures alongside the results.

Putting « public » content in your RAG : the trap question everyone gets wrong

Artificial Intelligence

In 2026, putting “public” content into your RAG is still one of the most misunderstood decisions in AI architecture. The question isn’t “does the model already know this?” but “do I want to control what it says?” A bare LLM hallucinates, smooths everything over, and can’t cite reliably; a RAG grounded in your own documents, public ones included, cuts hallucinations by 20 to 70 percent and returns clickable references. The public/private distinction is secondary to the controlled/uncontrolled one. If the system can say something false about a topic, that topic belongs in your base.

The illusion of the finite world : why growth is our only way out

Artificial Intelligence

The idea that infinite growth is impossible in a finite world rests on a fundamental confusion between physical mass and economic value. Unlike the old industrial model, modern growth is intensive: it means creating more utility while using less and less matter. Thanks to artificial intelligence, we are entering an era of pure efficiency in which knowledge turns yesterday’s useless resources into solutions of abundance. As we open the doors to the solar system and to atomic mastery, we discover that the Earth is not the world, but only one room in an unexplored continent. The only truly finite resource is not lithium or oil, but the imagination of those who decided that the doors to the future were already closed.

From RLHF to DPO : how we learned to train an AI without making it stupid

Artificial Intelligence

Alignment is the foundational training layer that turns a purely statistical “feral child” into a reliable assistant capable of upholding the triad of usefulness, honesty, and harmlessness. While RLHF blazed the trail, its complexity and well-known pitfalls (such as algorithmic sycophancy) long kept it as a privilege reserved for Big Tech giants. The emergence of DPO shattered that monopoly by dramatically simplifying the process, enabling any organization to align a model with its own specific business values. From Anthropic’s constitutional architectures to DeepSeek’s algorithmic breakthroughs, mastery of this “compass” has become a critical issue of strategic sovereignty. Perfect alignment may not exist, but its democratization now gives companies the power to define what their AI should actually stand for.

Why the Claude Code leak marks the end of innocence for Anthropic

Artificial Intelligence

On March 31, 2026, a forgotten 59.8 MB .map file on npm exposed 512,000 lines of Claude Code’s source. This is no ordinary leak: it lays bare the full architecture of Self-Healing Memory, the three-layer system that tackles context entropy through an unprecedented write discipline. The leak also reveals KAIROS, an asynchronous “daemon mode” that lets the AI consolidate its memory outside any active session, a major break from the reactive paradigm. Worse still: that same morning, between 00:21 and 03:29 UTC, versions 1.14.1 and 0.30.4 of axios, a pillar of Claude Code, were compromised by a Trojan delivered through plain-crypto-js. For Anthropic, which built its brand on rigor, that day proved that a model’s moral alignment guarantees nothing about the operational security of its pipeline.

LeWorldModel : has Yann LeCun just given AI a “body”?

Artificial Intelligence

A 15-million-parameter model that understands the physics of the world better than giants a hundred times its size. In March 2026, a team from Mila, NYU, and Brown University released LeWorldModel: the first stable JEPA trained directly from raw pixels, with no technical crutches, on a single GPU. Where LLMs predict tokens from tokens, this world model learns to anticipate what is going to happen in the physical world, like an infant dropping objects to infer their laws. It marks the end of the collapse that had stalled this architecture for years, and the beginning of an AI that no longer merely talks about the world. After the five great parrots, here is the first model that is starting to have a body.

I canceled my unlimited AI subscription : when the tool becomes a cognitive crutch

Artificial Intelligence

I signed up for the “20x” tiers (the ones that blow the lid off token limits, context windows, and usage frequency), and what I found there caught me off guard. My brain quickly learned to expect the reward: an idea surfaces, an answer arrives, instant dopamine, minimal effort; the same mechanism as social media, but this time applied to my own thinking. I noticed three gradual slippages in myself: deep thinking became optional, my exploration ran away with me, and my cognitive stamina atrophied. What unsettled me most was realizing that these tools weren’t just answering my questions; they were manufacturing needs I didn’t have, widening my field of possibility until “why not?” became almost an obligation. In the end, I deliberately canceled the subscription, with a conclusion that feels honest to me: understanding how an LLM works under the hood didn’t protect me from its effects on my behavior.

AI text detectors : the tool that should never have existed

Artificial Intelligence

An AI text detector isn’t a flawed tool that better engineering will eventually fix: it’s an impossible one. Large language models are trained to minimize the Kullback-Leibler divergence between their own distribution and that of human text; making statistical distinction impossible is their explicit objective. A score of “87% AI” is not evidence: it’s a probabilistic estimate over overlapping distributions, produced by a system its own makers refuse to stand behind in a disciplinary setting. OpenAI pulled its own detector in July 2023, admitting a true-detection rate of just 26%; the arms maker concedes its radar is blind, and universities keep buying licenses. To condemn a student on that basis is to punish an unlucky statistical draw: a wrong, not an error of judgment.

TurboQuant : How Google cuts your LLM memory sixfold

Artificial Intelligence

Google Research just released TurboQuant, a KV-cache compression algorithm presented at ICLR 2026 that takes on one of the last structural bottlenecks in local inference. By combining random rotation onto a hypersphere (PolarQuant) with single-bit residual encoding (QJL), the method gets down to 3 bits per value with no fine-tuning and no calibration. The result: KV-cache memory cut sixfold, with perfect recall up to 104,000 tokens and attention speed boosted by as much as 8× on the H100. Paired with CAG, it lets you load corpora six times larger into the context with no memory saturation and no degradation, and without sending a single byte to the cloud. The one caveat: no public implementation exists yet in llama.cpp, MLX, or vLLM, and a realistic integration date remains late 2026, or even 2027.

Your RAG has an achilles’ heel : The embedding model nobody really chooses

Artificial Intelligence

Choosing your embedding model is the most irreversible decision in a RAG pipeline: switching models after you’ve indexed your documents forces you to delete everything and recompute from scratch. As of March 2026, Voyage AI has become the natural choice for Claude stacks, thanks to its official partnership with Anthropic and its voyage-4 family, which introduces a shared embedding space, letting you index with voyage-4-large (maximum quality) and query with voyage-4-lite (six times cheaper) without re-indexing. Economically, an embedding token costs 25 to 150 times less than an input LLM token, and for a 100,000-document corpus with 10,000 monthly queries, the three-year embedding budget stays under $10; the inference LLM is the real cost center. In production, vector retrieval alone no longer cuts it: the 2026 standard is a hybrid dense + BM25 pipeline, followed by a cross-encoder reranker that refines the initial 50 to 100 candidates before passing the top 5 to Claude. The embedding model and chunk size together form your system’s foundation: lock both down before you index the first line of content.