GitHub Copilot: The bill comes due

There are moments when you use a technology and think: this is too good to last. Enjoy it while it lasts, before it is gone. The price for this performance is too fair. There had to be a catch somewhere. I had been thinking about this for a while when using GitHub Copilot. Now I got my answer.


Somehow I always had a feeling this moment would come. I had just hoped it would come later.

How it began

I still remember clearly the first time I seriously got into agentic coding. That was with Blackbox, a tool that showed me what becomes possible when you use AI not just as autocomplete but as a genuine development partner. I started improving code, delegating tasks, handing over context, questioning results, and refining with assistance. It worked. Quite well, actually. Then I switched to GitHub Copilot. And was convinced. I deposited my first budget and carefully monitored costs, as one does with new technologies.

It is important to understand: I am a UX designer. Development was never my core task. Even JavaScript development is hard for me. My domain is the "colorful things" in software — the look, how using it feels, the impression that remains when you are done. Complex backend development is not part of that, even though I have plenty of ideas in that direction. But with AI code assistance, my toolset expanded to what felt like infinity. Suddenly I too was able to build things I could otherwise only have realized with a development team. At a much, much higher price.

What emerged over the past months as an experiment was CrucibleMark. A modular LLM benchmark framework that I built module by module from concept to completion. Data architecture, scoring logic, export pipeline, political compass, ToolUse benchmark. Things I could never have realized at this depth and speed on my own.

And the GitHub budget grew — not despite better usage, but because of it. First 50 dollars, then 60 dollars, and eventually 100 dollars a month. Because Copilot developed rapidly during this time. From a simple single-thread development chat to a full agentic workflow with multiple subagents, tool integration, and autonomous work steps. What I got in return was an assistant that grew with me. And I was willing to pay for it — also because compared to a development team, the monthly subscription still won by a wide margin.

A trend I saw coming

But I was not oblivious. The market had already been sending signals. Anthropic was the first provider where I observed billing becoming more complicated. Agentic token consumption was suddenly no longer part of the standard subscription. Tool use, agent workflows, intensive API usage — all things that barely register during normal chatting but quickly become the dominant cost factor when seriously developing with AI. Claude was good, but also the most expensive. What had previously seemed included was suddenly billed separately. And that was noticeable.

Then OpenAI. In my own benchmark data it was measurable: the cost of a GPT-mini-4.5 benchmark run was three times that of a GPT-mini-4.1 benchmark. Same task, a significantly more expensive result.

You constantly read in the media about the immense infrastructure costs on one side and the incredible investments on the other. Keywords: Nvidia and others, keeping this development running with investments on the scale of an industrial nation's annual budget. I have been wondering for a while: how is this supposed to pay for itself? Whether a few chat subscriptions can really sustain that? And looking at the competition between OpenAI and Anthropic, it is clear that pressure is growing from all sides.

GitHub Copilot was my island for a long time. Affordable, powerful, seamlessly integrated. Here I could use Claude Sonnet assistance at a genuinely good price. Which raises the question of how Microsoft could sustain that. Either the other providers were overcharging us users, or Microsoft was subsidizing its offering. Now I know: they were subsidizing. And from June 2026, the good times are over.

The announcement

Then this message arrived from GitHub.

Starting June 1, 2026, Copilot is switching to usage-based billing. PRUs will be replaced by AI Credits.

Sounds technical at first, almost harmless. Then I looked at the dashboard they provided, which let me see what my April usage behavior would have cost under the new model. A fair approach — and a sobering realization: an increase from just under 100 dollars to 850 dollars. Included in the plan: 39 dollars in monthly subscription costs.

GitHub Copilot Billing Dashboard April 2026: Current Billing (PRUs) $97.72 versus usage-based billing (AICs) $808.32 – cost increase of around 730 percent with identical usage behavior
The GitHub dashboard shows: with identical usage behavior, April 2026 under the new billing model would have cost $808.32 instead of $97.72.

That is eight times as much. Anyone who has been spending 80 to 100 dollars a month and maintains the same usage behavior will pay 800 to 900 dollars going forward. Anyone who sticks to their previous budget will get only a fraction of the previous output in return. Honestly, I cannot imagine how I am supposed to keep working effectively under these conditions.

Is this the end?

No. But it is the end of an era. Nobody could have anticipated two years ago how this journey would unfold. The cost explosion, the global demand, the speed at which an entire industry has restructured itself. This phase was like riding a bucking bull. Nobody really had control. The momentum swept the business models along with it.

And as with every new technology, every new startup, the same principle applies here: first comes customer acquisition. Then comes the bill. The AI providers have made all of us — myself included — dependent on their offerings with digital candy. They let us experience the advantages of the new way of working. At some point the habit sets in, and the price becomes less and less of an issue.

I no longer want to go back to where I was just a year ago. A workflow without AI assistance means for me today: slow progress, enormous energy expenditure, significantly slower development. And I am not an isolated case. How many companies have already restructured so deeply that going back is barely conceivable? Because the bill always comes due. And now it has.

The calculation flips

But where one door closes, another opens. Given the costs to be expected going forward, a topic I have been thinking about for a while moves back to the foreground: building my own local infrastructure that lets me use AI in a privacy-compliant, independent way, without monthly surprises.

Until now, this consideration always seemed premature. At 100 dollars a month for GitHub Copilot, I calculated: even over three years, a cloud subscription is cheaper than a high-performance local machine. The only real downside was privacy — the fact that my code and prompts flow off to the US. Unsatisfying, but tolerable.

But now the wind is shifting. Concretely, two options are on the table for me. The first: I replace my MacBook Pro with a 128 GB unified memory beast. This is no ordinary MacBook Pro anymore — it is a machine from hell in laptop form. With 128 GB of RAM, serious local models can be held entirely in memory without compromise. The price, at around 6,000 euros, is at the extreme end. Local, quiet, and portable everywhere.

The second option has interested me for a while: the ASUS Ascent GX10, a compact supercomputer for local infrastructure, with an Nvidia chip and 128 GB unified memory designed for local use. The price: around 4,000 euros. Small, unobtrusive, powerful. A device I do not yet own but have been thinking about since it was announced. Stationary, running on the home network, as a dedicated local AI station. Less mobile than a MacBook, but for everything I develop at my desk during the day, more than sufficient. And at the currently exploding cloud prices, such an investment suddenly pays for itself within a year. Maybe even faster.

Is this the beginning of the open-weight era?

What no longer worries me is the quality question. Just two years ago, "local" was synonymous with "worse." That is no longer true. In the CrucibleMark data it is measurable: models like GPT-OSS, Qwen 3, Kimi-K2, and Mistral derivatives reach a level in many task areas that can hold its own against commercial models. What was previously a compromise increasingly is not.

And so something interesting is happening. The price increases from the major providers deliver exactly the economic pressure that could drive local and open-weight alternatives into mainstream practice. Not out of idealism or open-source conviction. But out of economic calculation.

… because the bill always comes due

Nobody could have anticipated how this journey would unfold. With commercial cloud models, a new technological era was established. This technology has shown what will be possible going forward, has accustomed and bound developers to AI assistance, and has made all of us more productive than we could ever have imagined. And now, when nobody can do without it anymore, the bill arrives.

But it seems as though open-weight models have become good enough at precisely this moment that the question is no longer: Can I work locally? But rather: Why not now?

I think I will find out. I already have the benchmarks for it.