Open source
Open weights won a quiet war in 2026. Everything we have written about open models, local AI, and the freedom trade-off.
-
A calm guide to running a decent AI model on your own laptop
Running your own AI model used to be a project for people with a spare graphics card and a weekend to lose. In 2026 it is closer to installing an app. Here is the calm, no-hype version of how to get a genuinely useful model running on your own machine, and honest expectations about what it will and will not do.
-
DeepSeek V4 is cheap, open, and quietly excellent
DeepSeek V4 arrived in April with the same trick DeepSeek always pulls: numbers that should cost a fortune, at prices that do not. Top-tier coding scores, a million-token context, an MIT license, and output priced under a dollar per million tokens. The West still has not quite made peace with it.
-
Gemma 4 12B review: how good is a model that runs on your laptop now?
Gemma 4 12B is not going to top any leaderboard, and reviewing it against the frontier would miss the point entirely. The question that matters is different: how much useful AI can you now run entirely on your own laptop, with the internet unplugged? The answer, it turns out, is a surprising amount.
-
GLM-5.2 is a 753-billion-parameter open model with an MIT license. That is a big deal.
Every month brings a new open model, and most are footnotes. GLM-5.2 is not. Z.ai released a 753-billion-parameter model, with a one-million-token context window, under a plain MIT license. That last part is what makes it matter.
-
Llama 4 Scout review: the giant context window, tested honestly
Llama 4 Scout is easy to review badly. You quote the ten-million-token context window, call it revolutionary, and move on. Having actually run it, the honest review is more interesting, and more useful: Scout is a very good open model whose best feature is not the one on the box.
-
Meta Llama 4 ships a 10-million-token context window. Do you actually need it?
Meta Llama 4 family landed with a headline number that is hard to ignore: Scout, the smaller variant, offers a ten-million-token context window. That is the largest of any model, open or closed. It is also a feature most people quoting it will never actually use.
-
NVIDIA Nemotron 3 Ultra quietly bets against the pure Transformer
Most model releases are the same recipe made a little bigger. NVIDIA Nemotron 3 Ultra is a bit more interesting, because it quietly does something different under the hood. It is a 550-billion-parameter model built on a hybrid of Mamba and Transformer architectures, and that architectural choice is the actual story, not the benchmark line.
-
Open weights or closed model: the trade-off that never goes away
The open-versus-closed argument gets treated like a team you join, complete with jerseys. It is not. It is a trade-off you make one job at a time, and the same person can sensibly land on opposite answers for two projects in the same week. What follows is a decision guide, not a leaderboard, because the leaderboard changes monthly and the trade-off underneath it does not.
What each side actually gives you
Open-weights models, the Llamas and DeepSeeks and Qwens of the world, hand you the actual model. You can download it, run it on your own hardware or a rented box, look at what it does, fine-tune it on your data, and keep the whole thing behind your own firewall. Nobody meters your calls. Nobody can change the model out from under you or deprecate it next quarter. If your data cannot legally or comfortably leave your walls, this is often the only real option.
Closed models, reached through an API, hand you a result. You send text, you get text back, and someone else owns the running of it. In exchange for giving up control you get the current frontier of quality, no infrastructure to babysit, and a model that quietly improves without you lifting a finger. For a lot of teams that is the entire pitch and it is a good one: you want the answer, not a second job running GPUs.
The costs hide in different places
People compare these on the wrong axis. They look at the per-token price of a closed API, see a number bigger than zero, and conclude that self-hosting an open model is cheaper. Sometimes. The closed price includes the hardware, the scaling, the uptime, and the salaries of people who keep it running. Open weights are free to download and very much not free to operate. You are now buying or renting GPUs, and someone on your team owns keeping the thing up at 3 a.m.
The honest version is about volume and steadiness. Spiky, low, or unpredictable traffic almost always favors the closed API, because you pay only for what you use and nothing when you are idle. Heavy, steady, round-the-clock traffic is where owning the hardware can win, because a machine you have already paid for does not care how many calls you push through it. The crossover is real, but it sits much further out than the sticker-price comparison suggests, and it moves every time GPU rental prices or token prices shift.
How to actually choose
Skip the identity and answer a few blunt questions about the specific job.
- Where is the data allowed to go? If it legally cannot leave your infrastructure, that decides it before any quality debate starts. Open weights, self-hosted, done.
- Do you need the absolute top of the quality range? For the hardest reasoning, the newest capabilities, the widest language coverage, the closed frontier models still tend to lead, and the gap is often worth paying for. For a well-scoped task, a mid-size open model may clear the bar with room to spare and cost far less to run.
- Does the model changing under you break things? If you have tuned prompts against exact behavior and cannot afford a silent update, a weights file you pin and control has an edge a hosted endpoint cannot match.
- Do you have people to run it? Serving a model in production is real, ongoing engineering. If that team does not exist, the API is not a compromise, it is the sane choice.
The gap keeps closing, the choice does not
The genuinely good news is that open weights have gotten shockingly close to the closed frontier on many everyday tasks. For summarizing, extraction, classification, ordinary chat, drafting, the practical difference is often small enough not to matter, and it keeps shrinking. That is a real shift and worth being cheerful about.
It does not, however, dissolve the trade-off. Control and privacy and the ops burden that comes with them sit on one side. Convenience and frontier quality and someone else's pager sit on the other. That tension is structural. It will still be here when today's model names are forgotten. The people who get the most out of AI are not the ones who picked a side and defended it. They are the ones who ask, for this specific job, which set of headaches they would rather have, and then pick accordingly.
-
Open weights won a quiet war in 2026
There was no dramatic moment, no single announcement, no headline that said it plainly. But somewhere in 2026, without a parade, open-weight models stopped being the scrappy underdog and became the default sensible choice for a huge slice of real work. It was a quiet war, and the open side won more of it than anyone expected.
-
The best cheap LLM in 2026 is probably not the one you think
Everyone reviews the frontier. Almost nobody carefully reviews the budget shelf, which is a shame, because that is where most real work should actually run. I spent time putting the cheap models through the same ordinary tasks I use every day. The winner was not the one I expected, and the losers were instructive.
-
The price of frontier AI just fell off a cliff
Quietly, without a single headline capturing it, the cost of using capable AI collapsed this year. Not dropped. Collapsed. The kind of model that cost a small fortune to run at scale two years ago now costs cents per million tokens. If your business plan assumed AI would stay expensive, it is time to redo the math.
-
This month in AI, sorted by what will still matter in a year
June and early July gave us a model release almost every day, which is exactly why you should not try to follow all of them. Most were incremental. A few will still matter next summer. Here is the month sorted the only way that is useful: by how long it will stay relevant.