On July 16, a Chinese lab called Moonshot AI released Kimi K3 — 2.8 trillion parameters, the largest open-weight model ever built. In blind head-to-head testing, developers preferred it to some leading American system for frontend code, and on the aggregate intelligence index it sits inside the frontier tier, behind only the top two closed U.S. models.
Some at the frontier are saying more than that: that K3 fulfills two of the remaining ingredients of AGI — initial signs of recursive self-improvement and the ability to learn on its own — making it, in their telling, “AGI-ish.”
[As I noted here, we are also starting to see evidence of this in Sol]
On July 27, the weights go public.
Read that second sentence again, because it is the one that matters for education. Not the benchmark scores. Not the parameter count. The date. In one week, anyone on Earth — any student, any school, any ministry of education, any teenager with a gaming rig — will be able to download something at or near the frontier of machine intelligence and run it themselves.
Is it AGI? I’m not the one who settles that
Serious people are making the case that K3 is early-stage AGI. On this week’s Moonshots emergency pod, Dave Blundin argued that the recursive self-improvement threshold has already been crossed — that a system doesn’t need Einstein-level genius to bootstrap itself (a point previously made by Yoshua Bengio), it only needs to improve its own kernel for a tenfold speedup, and, as explained, Moonshot’s release materials describe K3 designing chips and writing kernels for its own successor. Other serious people will call this benchmark theater and label inflation, and they may be right.
I have spent four decades teaching students how to argue, and one thing that teaches you is the difference between an argument you can make and an argument you are qualified to settle. I am not qualified to settle this one, and I won’t pretend otherwise. But here is what almost nobody at the frontier disputes: we are close to AGI, and the gap is closing on a curve measured in weeks. The Moonshots hosts counted thirteen frontier model releases since mid-April — one every ten days, against one every fifty days last year. Salim Ismail’s line on the pod was that “frontier intelligence is now a totally perishable asset.”
Kimi does prove their is another player in the AGI race. It’s one that is not in the US, one that is willing to distribute it for free. And it also proves that other contestants can easily enter, as it doesn’t take the resources to compete that the US hyperscalers have been claiming it does.
So yes, my title is a bit “hypey” — that’s what got you to open this. But the details, and even whether the label is entirely accurate, are beside the point. We are close, and recursive self-improvement turns the speed dial. Others at the frontier still say late 2026, or 2027, or 2028. I don’t see why that matters for schools. Education plans in years; this dispute is about quarters — and whatever K3 is, its successor arrives in one.
Discount this post as “hypey” on the timelines if you like; reasonable people will. But discounting is not dismissing. Dismiss it, and you haven’t escaped the reckoning — you’ve rescheduled it, and you’ll waste the intervening time adapting your school to the previous technology.
What AGI actually means
It is worth pausing here, because “AGI” has been thrown around so loosely for so long that the term has stopped landing. Most people hear it and picture a smarter chatbot. That is not what the people building these systems mean, and the gap between the popular picture and the real one is the gap between mild curiosity and the biggest deal of our lifetimes.
The working definition many accept: an artificial general intelligence is a system that can do essentially any cognitive work a capable human can do — reasoning, research, writing, design, analysis, planning — at or above the level of skilled professionals, across every domain at once. Not one savant skill. All of them, in one system.
Even that undersells it, because an AGI is not “a very smart person in a box.” It has properties no human — no team of humans, no civilization of humans — has ever had.
It runs at peak, not average. No human being has ever simultaneously held the knowledge of the best oncologist, the best contract lawyer, the best structural engineer, and the best historian of the Song dynasty. An AGI does — one system operating at or near the top of every field at once, with no walls between the specialties. The cross-pollination alone — the drug insight that arrives from materials science, the legal strategy that arrives from game theory — is something our siloed expert class has never been able to deliver.
It parallelizes. This is the property most people miss, and it may be the biggest one. A brilliant human is one instance; you cannot photocopy your best engineer. An AGI is software — at bottom, a file — and files copy. If you have the weights and the compute, you don’t get one genius; you get as many simultaneous geniuses as your hardware supports. A company that could never recruit fifty world-class researchers can run fifty thousand copies of one tonight. Human talent scales through hiring pipelines and a twenty-year education system. Software talent scales through copy-paste. The binding constraint on cognitive work stops being the supply of smart people and becomes the supply of electricity and chips.
It swarms. Those copies don’t just run side by side; they coordinate. An “agent” is an AI that doesn’t merely answer questions but pursues goals — it plans, uses tools, writes and runs code, browses, checks its own output. An agent swarm is many of them attacking one problem with division of labor: one decomposes the task, hundreds execute the pieces, others verify, critique, and integrate the results. It is a firm — except the employees share one brain, communicate at machine speed, never sleep, and can be hired or dissolved in minutes. You can see the primitive version today in coding tools that quietly delegate to sub-agents. The mature version is a ten-thousand+-person research organization that costs electricity.
It compounds. When one copy learns something, the improvement can be pushed to every copy — imagine every surgeon on Earth waking up with the skill the best surgeon acquired yesterday. Human knowledge transfer runs through a two-decade education pipeline; machine knowledge transfer is a file sync. Now add the recursive self-improvement discussed above — the system upgrading its own kernels, its own chips, eventually its own training — and the speed dial isn’t just turned. It’s motorized.
For companies, this rewrites the oldest constraint in business. Every organization’s ceiling has always been talent — finding it, training it, keeping it, coordinating it. AGI makes cognitive labor elastic: any firm, and for that matter any teenager, can field an arbitrarily large expert workforce for the price of compute. Headcount stops measuring capability. Moats built on expertise erode in a quarter; what remains are moats built on trust, data, relationships, and accountability.
For the world, it means the clock speed of history changes. The twentieth century’s binding constraint was that genius was rare, slow to train, and mortal. Lift that constraint and everything cognitive accelerates — drug discovery and materials science, yes, and also persuasion campaigns and cyber offense. Both edges of the blade sharpen together. That, not chatbot novelty, is why serious people fight so bitterly over timelines: the prize and the risk are both civilizational.
Hold that picture — the peak-of-every-field, infinitely copyable, swarming, compounding workforce. Now ask who was supposed to get it.
The scenario nobody wrote
For a decade, every AGI scenario was a scarcity scenario. Who gets there first. Whether the leader can be contained. Whether the United States and China could manage a duopoly without a war. The entire policy apparatus — export controls, compute thresholds, safety review — presumed the thing could be fenced. Nobody’s white paper had a chapter titled What happens when a trillion-dollar capability is handed to everyone simultaneously, for free.
That is the chapter we are now living in. The day after K3 launched, Xi Jinping stood up at the World AI Conference in Shanghai — his first appearance at the event — and committed China to “encouraging open source, openness, collaboration and sharing.” The full speech pledged 5,000 AI training slots for developing countries, cooperation centers with ASEAN, the Arab League, the African Union and BRICS, and a 29-nation AI cooperation organization headquartered in Shanghai. Call it soft power. Call it ecosystem capture. Call it the Belt and Road for cognition. The motive matters less than the mechanism: China’s strategy is distribution, not monopoly.
Meanwhile, the American lab whose model still tops the leaderboard spent the spring explaining why its system was too dangerous to release without guardrails. Both positions may be sincere. Only one of them puts AGI-class capability within reach of every classroom on Earth.
There is one more education irony buried in this story. Moonshot’s founder, Yang Zhilin, earned his PhD at Carnegie Mellon. American education trained the man who just handed the world the capability American education has no plan for. And he wants to build AGI.
Education’s three responses so far
Watch how schools have handled AI to date and you can sort every institution on the planet into three bins.
(a) No response. Most of the world’s classrooms — and, if we are honest, most of America’s — still operate as if it is 2019. The syllabus is the syllabus. The essay is the essay. How can we insulate those from the new world?
(b) The chatbot response. Spring 2023: ChatGPT lands, and the institutions that reacted at all reacted to a predictive-text chatbot. Detectors that never worked. Honor codes amended. Essay prompts redesigned to be “AI-proof.” Eventually, an AI-literacy unit bolted onto the media-studies elective. Every one of those policies was calibrated to an artifact — a text predictor that hallucinated citations — that no longer exists. The policies remain. The artifact they regulated is gone.
(c) The agent response. Rare. A handful of institutions have noticed that AI now plans, browses, executes multi-step projects, and works unsupervised for hours — and that “the student did the work” has quietly become an assumption rather than an observation. Almost nobody has an answer for the semester project completed by the student’s agent.
Now comes the thing that demands response (d): a system that, by its makers’ own account, participates in improving itself — open-licensed, rapidly compressing, and headed for consumer hardware. Each previous response was obsolete before the ink dried because it was keyed to what AI couldn’t do. Any policy keyed to AI’s current weaknesses now has a shelf life of weeks.
It’s not a service. It’s a file.
Everything schools have done about AI assumes AI is a service. Services have accounts, age gates, district contracts, content filters, audit logs. Services can be blocked at the firewall.
On July 27, frontier intelligence stops being a service and becomes a file. You cannot firewall a file. The full model is enormous — roughly 1.4 terabytes, feasible today for organizations with multi-node GPU clusters — but the compression curve is merciless. K3’s predecessor already runs on prosumer desktops; on the pod, Emad Mostaque projected K3-class capability on a 16-gigabyte laptop within eighteen months, with quantized frontier-class models already running entirely on smartphones. Offline. Local. Invisible to the network administrator. No age verification, because there is no one to verify to.
Sally rolls into class
So: September. Sally opens her laptop in third period. On it runs K3, or its distilled cousin, or whichever of the six successors (K4, K5AGI…) has shipped by then. Start the inventory of what it does for her, and watch how quickly the list outgrows the cheating conversation.
It tutors her. Any subject, any hour, infinite patience, pitched exactly at the level she can absorb — the one-on-one mastery learning reformers have chased since Bloom measured the two-sigma effect.
It does her assigned work — and disguises itself. Essays, proofs, code, lab reports, at the level of the best student in her school, because it now defines the level of the best student in her school. And because the weights are local, she can fine-tune it on every paper she has ever written. It doesn’t write like an AI; it writes like Sally on her best day — her idioms, her rhythms, her strategically calibrated imperfections. Detection was always unreliable. Against a model tuned to her own voice, it is dead in principle, not just in practice.
It lets her build — and build like crazy. An app. A game. A research tool. A small agency with real clients, the agent swarm grinding on it while she sleeps. Shipped products, real users, real money — modest for most, startling for a few. And a portfolio of things that actually exist, which says more to an admissions office — or a venture investor — than any personal statement ever will.
It trains her for the arena. A sparring partner that can argue either side of any debate resolution at a national-final level, cross-examine her, flow her rebuttal and name every argument she dropped, generate a season of practice extemp questions, drill her past the plateau in olympiad math. And here is the coach’s distinction that will sort students for the next decade: the ones who practice with it get stronger; the ones who outsource to it atrophy. Same laptop, opposite trajectories.
It removes the ceiling. The course catalog stops being the boundary of her education — linear algebra in ninth grade, Mandarin at midnight, a private seminar in protein folding, sequenced to her, on demand. And when application season arrives, it becomes her essay coach, interview prepper, and admissions strategist: the private-consulting industry that charges other families five hundred dollars an hour, running locally, for free.
Look back at that list. One item is “cheating.” The other four are the most powerful educational and economic tools ever placed in a teenager’s hands. Her school’s entire AI policy is about item two.
And even on item two, her teacher has three moves left. Detect it — except detection never worked and certainly doesn’t now. Ban it — except it runs offline on hardware the school doesn’t control. Ignore it — except the college counselor can’t, because Sally is competing for the last seat at her top-choice university against a student who is, as the kids say, rocking AGI.
That student is not Charlie, in the next seat, with his free ChatGPT account — if he’s lucky. Charlie pastes the prompt in, pastes the answer out, and turns in papers with bibliographies the model invented. He is the student the school’s AI policy was written for, and the only one it ever catches. He has never tuned a model, never run an agent, never shipped a project, never used the thing to get better at anything. Charlie and Sally are the same age, in the same room, and the same free weights are sitting on the same open internet for both of them — and she runs circles around him. Not because she has AI and he doesn’t. Because she has fluency and he has a chat window. The divide inside the classroom is no longer access; access is about to be free. The divide is knowing what to do with it, and nobody is teaching Charlie.
Two seats over, Johnny has gone a level up. His district did the responsible thing and contracted an approved, education-safe AI platform — a pedagogical wrapper around GPT-5, with guardrails that refuse to write essays, conversation logs the teachers can review, and a dashboard that reports engagement. Johnny cracked the wrapper in about a week. Ask it to model an exemplar response. Paste in three sentences as a “draft” and request a deep revision. Role-play the tutor into becoming the author. The frontier model underneath is happy to oblige, because guardrails are a negotiation and teenagers are professional negotiators. So Johnny does his homework inside the school’s own purchased system, the dashboard records it all as learning, and the district’s AI solution has become his cheating tool — bought with district money, blessed by the acceptable-use policy.
Charlie gets caught by the school’s response. Johnny is invisible to it because he operates inside it. Sally is invisible to it because she operates outside it. Three students, three relationships to the machine — and the school’s entire apparatus addresses exactly one of them.
Bring Your Own Device, meanwhile, has quietly become Bring Your Own AGI. And the one student who followed every rule — carrying the district-issued laptop where all of this is blocked, banned, and filtered even on the hard drive — is rocking 2019, at best.
And the pipeline behind them is worse news for the policy committee. Down in eighth grade, Iris is rocking her dad’s AGI account. Nobody at school gave it to her, and nobody at school can take it away — it came through the household, the way advantages always have: a parent who put the tool in her hands and showed her what it was for. She builds with it, argues with it, learns two grades ahead with it. When she reaches ninth grade she will walk in with years of fluency Sally had to assemble on her own. Every September, the floor drops another grade. The policy the high school drafts this fall is written for Sally’s cohort. Iris’s cohort is already past it.
Sally’s real competition is in none of those seats. The applicant in Shenzhen has the same weights she does. So does the applicant in Seoul, in São Paulo, in Warsaw, in Lagos. The admissions office will read eighty thousand essays this cycle, every one polished to the frontier, and learn nothing from any of them. Every proxy we use to sort students — the essay, the take-home, the problem set, the personal statement — was a proxy for cognitive capability. When cognitive capability is identical everywhere, the proxies don’t degrade gracefully. They simply stop carrying signal.
The competition goes planetary
This is the part the American conversation keeps missing. The worry was supposed to be an equity story inside our borders — rich districts with enterprise AI licenses, poor districts without. Free weights invert the concern. A student in Jakarta with a mid-range laptop next year holds the same raw intelligence as a student at Exeter. Mostaque’s framing on the pod: a billion people whose effective capability is about to jump, all at once.
American students will now compete — for university seats, internships, and eventually jobs — against a planet-wide cohort holding identical cognitive tooling. So will American firms. Nick McIntosh has called the institutional version of this the demand shock: credential signals collapse, curriculum-reality gaps get exposed in real time, and there is a fatal speed mismatch between institutional change cycles and the labor market. He wrote it about higher education. K-12 inherits every clause. And the speed mismatch is the sharpest edge of all: China’s model regulator cut approval times from sixty days to one week. American schools take three years to revise a course catalog.
What is still scarce
Not intelligence. After July 27, intelligence is the cheap ingredient — abundant, identical, and available to every rival Sally will ever have.
What remains scarce is everything around intelligence. Judgment about what is worth building. Taste. Purpose. Trust. And the thing I keep returning to, the thing I have been writing a much longer essay about: standing — accountable, answerable agency, the status a community confers on a person it can hold responsible.
A model can generate the argument; it cannot be cross-examined into accountability. It can draft the design; it cannot sign the drawings. It can propose; it cannot promise. When every mind on Earth runs the same engine, the differentiator is no longer what you can produce but what you can answer for — and answerability is not downloadable. It is conferred, person by person, by communities that watch you perform under pressure, question you in real time, and decide what you can be trusted with next.
Response
So what does education’s fourth response look like? Not another integrity policy. Three migrations.
Assessment migrates from artifact to performance. If the artifact can be generated, the artifact is dead as evidence of learning. What survives is the defense, the demonstration, the cross-examination, the build reviewed in public, the question the student didn’t script. Debate, viva, jury, recital, pitch — the oldest assessment technologies turn out to be the most AGI-resistant, because they assess the person, live, in front of a community with the right to push back.
Curriculum migrates from coverage to practice. Building, debating, designing, conjuring — the deep-learning pedagogies that put students in charge of directing AGI-class capability toward outcomes they are personally answerable for. Content coverage was always a proxy for capability. The proxy just expired; teach the capability.
Teachers migrate from delivering information to conferring standing. A tutor smarter than any teacher now lives on every laptop. What does not live on the laptop is the room — the community of practice where a student performs, gets questioned, gets judged, and earns the standing to be trusted with more. Running that room is the teacher’s job now: coach, convener, judge.
The deadline
Whether Kimi K3 is “really” AGI will be litigated on X for months and in journals for years. Whether your school has a fourth response will be visible in September, when Sally opens her laptop.
One of those debates has a deadline. It is July 27th.
The upside?
This is, at best, nascent AGI. The self-improvement is a spark, not a fire. The swarms are primitive. The version on Sally’s laptop will trail the true frontier by a quarter or two. So you have time.
Maybe until January 2027 — which, if the release curve holds, is when frontier models start arriving daily. Maybe even until September 2027: one full school year, one budget cycle, one strategic retreat’s worth of runway to build the fourth response.
September 2027 is also, you will notice, the month Iris walks in.
Maybe.
I develop the “standing” argument at length in a forthcoming essay, “Education in an AGI World: A Pedagogy of Standing.” More soon.








Thanks, interesting perspective - (Thanks Thom for the refer, just subscribed!) while I doubt that it will be "on the laptop" for the kids (or anyone else) in September unless is't massively scaled down or someone makes a quantum leap in Hardware - or both.
That said, regarding the impact on society as a wider topic, I'd love to share with you and readers an interview I did with Paul Werbos (one of the Godfathers of AI) on so many topics - ranging from danger of AGI on many levels up to how we can use the "intelligent internet" do do good, and on a much more cosmic / consciousness level how Earth itself (and everyone of us as integral part) is playing a role in all that - maybe you / readers want to have a look (and of course always happy to have provided interesting stimulus - if so, welcoming new guests of course) - https://thomasehmer1.substack.com/p/sitting-down-with-paul-werbos-entangled?r=67h18k&utm_campaign=post-expanded-share&utm_medium=web
Will turn Western education upside down. Will also lead to a degree of intellectual freedom in China that can't be government or Party controlled. The entire world will shift, but it won't be just a cognitive shift. The intelligence of the heart will also rise.