Right now we have a LOT of band aids. You want to optimize compute and thinking to a particular problem, sort of like we do. Yes you cannot perfectly predict this but you can do decently well and save a ton of tokens at the cost of this band aid being sort of leaky and gross.
But the larger problem is sound, and the answer is something jointly optimized (idk how they do the routing) but it’s hard to shoehorn it into the current paradigm.
There are at least two companies out there that can fold some laundry in a controlled environment already. At this point it's just a matter of categorizing fabrics and shapes and expanding that knowledge. Two years at the outside until it can fold 80% of your laundry.
Totally agree but the idea is this gives you a teleoperation environment that is truly on policy and not some artificial lab. The idea is that these robots, like those Amazon stores, are predominantly just controlled by actual humans.
Hey I totally agree I do not want a teleoperator looking into my house, it’s just so deliciously tempting to get in home on policy data. Not sure the reason why they are super interested in home environments vs business or public spaces.
This is basically a solved problem. Humans have latency, more than your Internet connection. We predict what's going to happen based on the shapes we see. Computer vision and LBMs do the same thing.
by your username I assume you are yourself one of the droids you're looking for lol.
But also I think people forget: this is not cut and dried. This is not simple. This is not just "these companies are evil and should be stopped". It is: there is a market pressure to do these things. "I enjoy the product and use it a lot" and "I am addicted" is blurry and market pressure is not going to recognize that limit because it does not care about human suffering unless that suffering meaningfully impacts the bottom line.
If these companies hit regulations that effectively cap their advertising revenue per user (i.e. the "addictiveness"), they are dead. That may be totally fine, and I'm sure majority of people would rejoice hearing this. But remember: advertising dollars are earned, especially at tech company scale, by the effectiveness of the targeting to get get more $ / DAU since DAU cannot grow beyond the Earth's population and that is the scale that these companies have already achieved.
If you cap advertising dollars, you cap advertising effectiveness. You cap the ability for small companies to connect quickly with prospective customers without being locked out because they have to spend too much to find them. Yes you also cap scammers and other nefarious actors too, but thats arguably a different issue. The impact of reducing advertising effectiveness is disproportionately concentrated on small business where cheap and effective advertising is so important.
My hope is that there is a way to do both and I don't have to be constantly horrified when I look at my screen time hours.
> by your username I assume you are yourself one of the droids you're looking for lol.
Aside from a 6 month contracting stint, I’ve never worked on data related to humans, only machines. No payment data, no behavioral data, no website data. The closest I’ve come is human-provided labels.
I’ve probably made only 50% of my total earning potential compared to Data Scientists / ML Engineers who work in the advertising and retail spaces (Google, Meta, Amazon), but at least I can look my children in the eyes and tell them that I didn’t sell my soul for a dollar.
> My hope is that there is a way to do both and I don't have to be constantly horrified when I look at my screen time hours.
Same, but without regulations there exists no incentive for any corporation to care about whether you’re addicted to your screen. And engagement, positive or negative, only makes them richer.
> This is not just "these companies are evil and should be stopped". It is: there is a market pressure to do these things.
They are evil tho, on top of there being market pressure. Their executives are incredibly comfortable with causing large scale harm, even when they are NOT forced at all.
> If you cap advertising dollars, you cap advertising effectiveness.
And that is OK. Not just ok, but actually good. Their effectivity and their evilness are closely related. Their effectivity and harm they cause are essentially the same thing.
> The impact of reducing advertising effectiveness is disproportionately concentrated on small business where cheap and effective advertising is so important.
Nice try, but no. Who benefited the most were the bad actors, conspiracy theories, weight loss programs specifically marketted to teens who posted about being insecured about their issues.
Marketing is just an arms race. It is ok to limit how manipulative and harmful it is and force companies to compete on something else.
> Marketing is just an arms race. It is ok to limit how manipulative and harmful it is and force companies to compete on something else.
I agree with this, and I think you are right about a lot of this being really just not market-relevant factors that are driving a lot of the downsides here. All of these should be regulated, and even the ones where regulating will have impact to markets and consumers.
And I agree -- effective advertising targeting necessitates "reward hacking" in that I can get you to buy something if I prey on your insecurities or otherwise manipulate you in a way that goes beyond "oh I know what watwut would love as a birthday gift".
But you really should be aware that effective advertising really does disproportionately benefit smaller businesses and this is generally a Good Thing. Really my only point here is that there is a baby in this bathwater and we should be careful.
This makes sense but I am not comfortable with the open source picture either. It depends on the use case and long term strategy. If you're handling customer support tickets or something, beyond some capability I can't imagine the ROI would be worth needing the absolute frontier. You are then in a wonderfully lovely place: absolutely full control over the model and inference stack (if you want).
But
- Plenty of businesses are stuck between a rock and a hard place: make 3p models load bearing, or sacrifice performance to the competitors that are willing to swallow that risk
- I just cannot imagine a viable equilibrium where OSS models compete on capabilities with the frontier without a hidden payer and shaky economics. Quant funds, some sovereign AI effort, cloud business, sanctioned distilled frontier models; all of these are demonstrably viable vehicles for OSS development. But the optimal position for these purposes is not to be the best, it's to be good enough (which I think is the position they find themselves in). I can imagine temporary points where open models pull ahead but not sustainably.
I agree there are many many use cases where the optimal choice is to reach for open weight models (I assume yours is one of them).
But the economics of frontier model development necessitates that these models are behind and thats a problem for other cases.
I am not sure I understand your arguments. You seem to be taking both positions. I suspect that is because there are a few confounding variables and diverse needs AI solves and that makes a cohesive analysis hard to reach. I have a clear view that cuts through the complexity, whether you find it useful or persuasive is another matter but I’ll lay it out.
1. All intelligence has value. A Harvard MBA or Stanford CS engineer have value. Haiku has value, Opus has value.
2. Sometimes it makes sense to buy, build or rent capabilities. But lack of ownership of something your business requires to operate is an existential risk which requires ownership.
3. Current OSS models are good enough for substantial automation and productivity gains with the right orchestration and harness. Especially with domain experts in the loop.
4. Frontier models while better, introduce existential risk; so they are a temptation that should be ignored if you care about your business being able to continue to exist.
Let’s invert this conclusion. Imagine you build a workflow around Fable 5 and the workflow going down would harm your business. Well the government has demonstrated ability and willingness to impose a multi day / week / month interruption in your business with no legally required notice, right to appeal or compensation. And further demonstrates a willingness to give preferential access to your competitors.
For these reasons, the only viable path is a model you have on servers you control and can replace if needed. That means cloud providers are fine if you have the model archived somewhere.
I agree with everything you're saying. And I would also agree if you can run your business adequately on OSS models you should absolutely do that for exactly the reasons you say. All I mean is: frontier models will always have their place and I don't see a way for OSS to ever change that picture. I also worry that the economics that makes meaningful open weights model development possible may break someday but I don't really see a mechanism for this at this point.
I don't actually have a use case in mind where you would _need_ or get a substantial competitive advantage of the absolute frontier in some sort of workflow or whatever load bearing DAG there is for your business. But if there are cases like this, that will suck.
Red hat did it for Linux. Someone needs to do the same for
AI. Make money on the implementation and maintenance not the api tokens. Tech companies have gotten fat and Happy on recurring revenue. But they are abusing the trust users have placed in them. I feel a back lash in a way I’ve never experienced.
I think AI labs are assuming that if they build and serve intelligence every one will want it. I think the last month has resulted in a lot of people saying “f u I would rather fail in life than serve you”. I personally went from feeling excited and optimistic about building amazing things on top of the models to hell bent on being self reliant and lots of my friends are doing the same.
I think the problem is whether you're for or against local models, the training strategies etc are roughly the same.
> I think the last month has resulted in a lot of people saying “f u I would rather fail in life than serve you”.
I agree a lot of people are saying this and doing this but this won't last long. People do not want to suffer or fall behind their peers. At some point not using AI for anything will not be a feasible option. Unless you operate from an incredibly uncommon and privileged financial position, you don't really even have a choice.
The thing the overlords are forgetting is they need customers and users.
You are right that the temptation to use the asymmetric advantage AI gives will be intense. But banks felt that temptation in 2008 in with cheap capital and 40x leverage. Some banks resisted. Which ones do you think survived in 2010? Business and technology moves in cycles. Optimize for longevity and let your competitors go out of business
Yes but in an equilibrium steady state, compute and data advantages are all you need to first order. China does not yet have a compute advantage. RL is indeed the magic sauce for coding agents but the bottleneck for how much progress you can make, for both the US and China, is compute. The US at least for the next few years has a clear advantage here.
Exactly -- the hope for US strategy is that you can slow them down a lot but not forever. That slowing them down is in itself enough to keep a strategic advantage over them both in terms of economic growth and offensive capabilities both in terms of cyber attacks, intelligence and things like drones, etc.
Definitely hasn’t backfired. Exporting would have just sped up their progress. Instead they had to get clever and lean into the bottleneck which for them, now, is compute efficiency. This is temporary and they’ll figure out competitive chip design and production but not for several more years. It’s incredibly hard to match the quality of NVidia and TSMC, as China has found throughout years of trying.
I’m worried if west gets to ahead the best thing for C is to destroy the factories to slow down the ai.
If west ai is too advanced can take over the world. So better go to war now on a same level playing field than later when you need to fight against a SGI
That is exactly the scenario that I believe and invest in. I peg trouble in Taiwan in the next 2 years at about 30%, which is waaaay higher than is priced into the market right now. If you think intel has gone up in price a lot now, it will absolutely skyrocket if TSMC fabs suddenly disappear. After adjusting to a domestic fab pipeline, we will have built up again an industry with a good talent pool (which we don't have now, Arizona TSMC fab needed to ship people in from Taiwan). At that point, why go back to a TSMC model? Hence we will have a booming domestic production pipeline, though still with complex international dependencies for various components.
Nope. If China were not banned from US-controlled chips, it would be importing chips at much higher prices, therefore getting less bang for its buck while strengthening competitors with its money in the process.
Instead, the US banned China from chips and lithography machines, giving China the legal excuse to start producing them domestically without violating WTO rules. Now China produces cheap chips and uses them with cheap electricity.
This was a dumb move by the US. Brought upon it by dumbf*ck aristocratic elites who grew up in isolated mansions and then received law degrees, with absolutely no understanding of technology and technology ecosystems. They thought they'd just make the rules and everybody would have to obey. It turns out in technology, they don't have to...
> Nope. If China were not banned from US-controlled chips, it would be importing chips at much higher prices, therefore getting less bang for its buck while strengthening competitors with its money in the process.
Nowhere near the value of having access to chips, at any cost. They have extremely deep pockets. They already pay 6x the cost per FLOP.
> Instead, the US banned China from chips and lithography machines, giving China the legal excuse to start producing them domestically without violating WTO rules. Now China produces cheap chips and uses them with cheap electricity.
You think without export restrictions China wouldn't be doing the exact same thing? China needs absolutely zero legal excuse. I mean sure they have compute available on grey market / domestically but at 6x the cost per FLOP. Access to NVIDIA chips would make it dramatically cheaper for them. Yes you get chip income but that is not even close to what you lose. The strategy is doing what it was always supposed to do: slow them down, bleed their resources to force them to spin their wheels catching up. China is doing a great job with this but they are fundamentally constrained by these export controls.
You are right that this greases the wheels, they are further along than they would have been without export restrictions, but they are still delayed even with the reduced friction. The alternative is that they move slightly slower _while having the same compute infrastructure available_ and at dramatically lower energy costs. That is a far worse position for the US to be in.
> This was a dumb move by the US. Brought upon it by dumbf-ck aristocratic elites who grew up in isolated mansions and then received law degrees, with absolutely no understanding of technology and technology ecosystems. They thought they'd just make the rules and everybody would have to obey. It turns out in technology, they don't have to...
I think this is too cynical. Neither one of us is in the room to actually observe the real decision making, but export restrictions as a strategy are not some "dumbf-ck aristocratic elite" thing. They are perfectly rational from a strategic standpoint and arguably doing what they're supposed to do.
> Nowhere near the value of having access to chips, at any cost. T
Disagree. Anything produced in the Eu or other US satellites is priced as high as possible. One could say that allowing China to buy those chips would have been a bigger drain on the Chinese AI budget.
> You think without export restrictions China wouldn't be doing the exact same thing?
I dont think it. China did not do it and stuck by the WTO rules after going through so many hoops to get in. Then the US decided that China was too competitive and started ignoring WTO rules.
> You are right that this greases the wheels, they are further along than they would have been without export restrictions, but they are still delayed even with the reduced friction
Disagree. Restrictions might have hampered them somewhat at the start, but setting up their own infrastructure as fast as China does and scaling it up as big and as fast as China does would eliminate any delay in a flash.
> I think this is too cynical. Neither one of us is in the room to actually observe the real decision making, but export restrictions as a strategy are not some "dumbf-ck aristocratic elite" thing. They are perfectly rational from a strategic standpoint and arguably doing what they're supposed to do.
Nope. The narcissism and aristocracy of the US elite is well documented all the way since the 60s since the biggest ongoing social study on the matter (from UCSC) literally mapped it:
First the US blocked China from buying NVIDIA's H100, but allowed NVIDIA to sell them a China-special nerfed H100, the H800
Then the US blocked the H800
Then the US realized that China was indeed accelerating their US independence, so does a U-turn and has now approved the H200 (more powerful than both the H100 and H800) for sale to China, on a case-by-case basis
However - and here is the real kicker - China themselves are now blocking H200 purchases since they want the acceleration towards Chinese homegrown solutions to continue, and now we have Chinese models being served on Huawei Ascend chips, with next generation Ascend 750 chips (using CXMT made memory) targetting training currently in testing.
Now we have Apple asking the US government for permission to buy memory from CXMT given the global shortage!
The Jensen argument that "oh we get them hooked on our technology and that will actually be better" is bullshit -- they would do the exact same thing they're doing now of building their own supply chains and compute power but be able to accelerate their progress in the meantime with US chips. It is worse in every way to sell them SOTA chips (or even previous generation chips).
Of course they have ways around this -- you can get black market GPUs and also API costs are SUPER cheap there -- they hack the subscription model, bundle a bunch of user accounts, and route API requests through them.
And yes they are getting to parity with US technology and will get there in a few years, they have decent chips but still not the quality of NVIDIA.
I don't know what you mean by the "quality" of NVIDIA. All the western AI accelerators, from NVIDIA, AMD, Google (TPU) and Amazon (Trainium) are made by TSMC, and their speed/density is only possible due to TSMC using EUV machines made by ASML (a dutch company).
Without access to ASML EUV machines, the Chinese will be stuck on older less-dense chip manufacturing nodes, but in terms of building a cluster this is just a cost/efficiency issue - it means you need more chips, more electricity!
You misunderstood my comment. My hypothesis is that it did _neither:_ it accelerated along an axis, and American SOTA/frontier is now laggard (efficiency).
Deepseek and Kimi are writing paper after paper with substantial architecture improvements for efficiency, because they can't just throw more hardware at the problem.
And China is now doing something on the hardware axis; which it may have never explored were it not for the sanctions.
> And China is now doing something on the hardware axis; which it may have never explored were it not for the sanctions.
Gaining parity on the semiconductor fab front has been official government policy as part of their Five Year Plan for at least the last decade, straight from the Politburo. They were always going to go down this path, and with AI playing front and center on their upcoming plan, there’s even more pressure.
There was never a possibility of them not exploring it.
There’s an ambiguity here: what is “this area”? If we’re talking about the semiconductor industry, they’ve been investing billions since the early nineties. The really big push came with the China Integrated Circuit Industry Investment Fund in 2014 at which point they were plowing billions into the industry across the public and private front. That one fund alone has $100 billion in assets under management at this point, and there are many other funds involved.
They’ve also invested in AI separately (before LLMs) in that time period but I’m less familiar with that sector.
Oh no Chinas chip manufacturing efforts have a long and rocky history, they would be pursuing a sovereign stack no matter what.
The tradeoff is worth it. They’re even publishing papers which blows me away — their efficiency gains quickly become incorporated into frontier models because they are open sourcing them. They would be aggressively pursuing the same chip pipeline strategy as they are today.
US is lagging in efficiency work because the ROI is better elsewhere for us. We have the same tier of talent, once the script flips so can the research.
Slow? The chip ban was just a few years ago. The result? China is more or less self-sufficient in chips, about to catch up with the latest generation. They even banned China from using lithography machines. The result? China is now producing lithography machines that only few in the world could produce.
Let's face it - all bans were dumb. They just gave China the legal (per WTO rules) justification to start producing everything domestically. The bans work as a reverse tariff, as a protectionist measure that actually protects your competitor. If China did those, others could bring China to court at the WTO. But the US did that, so nobody can sue China.
They are still a few years away from parity. It's not that they _cannot_ train frontier models its that it costs 6x as much roughly, and thats ignoring the headaches and extra hoops they need to gather that compute. That is the entire point of export controls. Slow them down, make R&D more painful. The whole picture -- Chinese research now being best in the world for compute-efficient AI development, China aggressively pursuing independence in compute capabilities -- are all perfectly predictable and do not invalidate the point that export controls are doing their job.
Regardless of any export controls (or really anything else) China would be aggressively ramping up chip production capabilities as fast as they possibly could. Export controls maybe help slightly here but it's worth it. China will succeed too, but it takes time. China has been pursuing this for decades.
> Let's face it - all bans were dumb.
Predictably, this is oversimplifying.
> They just gave China the legal (per WTO rules) justification to start producing everything domestically.
It just doesn't matter, reliance on chips from an adversary is an untenable place to be. They don't need any legal excuses.
> he bans work as a reverse tariff, as a protectionist measure that actually protects your competitor.
Again untrue -- tariffs are devastating economically but BOTH the US and China want independence from another, it's just a very painful cost economically to do the divorce. It's far less efficient and horrifies economists but geopolitics is multifaceted.
> If China did those, others could bring China to court at the WTO. But the US did that, so nobody can sue China.
Hmmmmm would this really work? Would this have any meaningful impact whatsoever on their strategic initiatives? I highly doubt that.
Idk if someone paying attention to how 2025 and 2026 have gone thinks that by 2028 we will be backing off of agenting coding that is wild. Like the other comment says: future models refactor the code of older models.
You can use LLMs heavily without ever actually "vibe coding". I do think to the degree "vibe coding" continues to exist there will always be work to do in turning some portion of vibe coded work into more robust production quality code. You can still use LLMs to do this you just have to maintain control over architectural choices.
Yea it’s hard for me to think of what the end state equilibrium is. A pile of vibe coded junk today is bad. You need humans. But we’ve made such a ridiculous amount of progress in such a short amount of time and, most importantly, this shows no sign of slowing down or plateauing. So will we hit a point where “vibe coding” is just all there is? Where human intervention is bad, just as hand tuning assembly is bad?
Is there a level of abstraction where human involvement will always be necessary? If so where?
I think it's not going to be a straight line upwards, it's going to get weirder. LLMs are still riffing on what we humans have done. If we stop giving them good examples because we aren't architecting anymore I don't really know what's going to happen. Probably some kind of meta architecture that emerges from lots of agents working in parallel but I worry it might have a strong house of cards theme to it.
if you were paying attention you would've noticed that between 2025 and 2026 the pricing of these things have somewhat changed. How does the extrapolation look with that?
This is just a misconception of how LLMs work and also what reasoning is.
“There cannot be any reasoning embedded in the model” a strong statement, what do you mean by reasoning because by any reasonable definition I’m aware of, they clearly are able to exhibit reasoning.
The fact that the pre training objective is next token loss has nothing to do with capabilities or their ability to reason. To be highly successful at next token prediction you NEED to reason. I’m quite confused here.
If you don’t mind actually taking a few more words to be more specific that would be helpful because what you’re saying doesn’t really make sense at all. You don’t need to trust that the reasoning traces are all faithful representation of an internal reasoning trace. Plenty of other ways to probe models (see anthropics work using circuit tracing).
What else is there to say? LLMs can at most regurgitate approximations of human reasoning steps in the limited forms in which they may be expressed in the training data or interpolations thereof. That's the core essence of what they are. There is no proper reasoning to be found.
"at most" is wrong. RL with verifiable rewards takes you beyond quality and skills represented in training data, I'm not aware of meaningful fundamental limits here if you scale compute enough even though right now it's highly sample inefficient.
Since you refuse to actually define what you consider to be reasoning let me at least put one out there: a system exhibits reasoning when an answer depends on nontrivial intermediate computation over the problem. If you find problems with this, fine, but just make an effort to contribute an alternative.
If you increase test time compute you get better performance. If the model was just "interpolating" this wouldn't really work would it? Models can do FrontierMath expert problems (unpublished, expert authored, peer reviewed math problems) that require an insane amount of compositional reasoning. If they were regurgitating training data, that wouldn't really work would it? Chain of thought, while not always faithful to internal computation, improves performance. If the models were just regurgitating information, it wouldn't work that well would it?
"regurgitating training data" is also of course misleading. Yea they can memorize parts of the training data, but they generalize very well.
There is the obvious limit that human text output is limited. To this you can add the specific testable training that pertains to code, but this degrades the weights for more general communication. Somehow the hype over the successes with coding in the last year or so made everyone forget the intrinsic limit posed by the exhaustion of real human text output, which is absolutely inescapable
> To this you can add the specific testable training that pertains to code, but this degrades the weights for more general communication.
Im not sure exactly what you’re saying here, is it that you trade off coding performance and performance on other tasks like communication? If so (correct if not) this isn’t true — generalization happens. Doing good on coding lifts all boats.
> Somehow the hype over the successes with coding in the last year or so made everyone forget the intrinsic limit posed by the exhaustion of real human text output, which is absolutely inescapable
You’re absolutely right that we’re quickly running out of human text data but that isn’t at all the limitation you think it is. No one has “forgotten” this — coding agent performance is primarily from reinforcement learning on synthetic data traces with verifiable rewards, though pretraining is still important.
Also don’t forget: there is a world of multimodal data (video, audio, 3D maps, etc) that is incredibly rich.
Reasoning includes things like proper use of logic. LLMs have been repeatedly shown to fail horribly at this.
They consistently fail at drawing basic logical conclusions because they cannot build a sufficiently abstract model of certain problems that allows them to grasp their true nature. In other words, the whole class of questions of the kind of "how many r's in strawberry" or "do I take the car to the car wash?" would be answered correctly and reliably.
> Reasoning includes things like proper use of logic. LLMs have been repeatedly shown to fail horribly at this.
That models cannot do ALL logic problems does not mean that they cannot properly use logic...they can write Lean-verified theorems. How is that not logic?
> They consistently fail at drawing basic logical conclusions because they cannot build a sufficiently abstract model of certain problems that allows them to grasp their true nature.
What does their "grasp[ing] their true nature" have anything to do with what they can do?
> In other words, the whole class of questions of the kind of "how many r's in strawberry" or "do I take the car to the car wash?" would be answered correctly and reliably.
Again, just because you have interesting failure modes or brittleness does not mean they do not reason.
This is exactly backwards. The brittleness is because they emulate reasoning without actually algorithmically performing it.
Add.: I pointed to this class of problems specifically because they require the ability to abstract in a way that the question itself does not immediately suggest. Math problems are different in that they are described in terms of art that are closely related to certain patterns of manipulation (that is, the paper texts tend to contain both in close proximity to one another).
For you, a system needs to reason perfectly and flawlessly, all the time? So humans do not reason? Humans don't have brittle failure modes?
> they require the ability to abstract in a way that the question itself does not immediately suggest
yes, yet there are multitudes of other measurements of the same kind where LLMs reason perfectly well and better in many cases than a human could.
> Math problems are different in that they are described in terms of art that are closely related to certain patterns of manipulation (that is, the paper texts tend to contain both in close proximity to one another).
Is your logic really that math problems are actually easier to answer without reasoning and just by blending together closely related papers? I would definitely suggest reading the literature a bit more on this topic.
Humans are not flawless, but they are much, much better at reasoning than LLMs are. LLMs can be made to fail quite reliably and easily because they cannot build proper manipulatable/predictive models. This is related to the point that Yann LeCunn makes when advocating for world models (for the physical world) with predictive power.
LLM output is a kind of dreaming but with the whole of past human text output as dream material. It turns out to be useful if you can direct the hallucination
But the larger problem is sound, and the answer is something jointly optimized (idk how they do the routing) but it’s hard to shoehorn it into the current paradigm.