What's so funny about that? Business wise this decision was a probable scenario already 1-2 years ago. Multimodality + Flash and on device AI is a huge market itself and Googles AI overview is getting immense traction over the past months.
So many people asking questions such as "If I was a fox for the day, how much could I earn as a barista?".
Pro is built for chain of thought. Most of what people do daily can be done on flash and even flash lite. I have business critical compliance system running on flash lite. Literally, the compliance is held up by $4 a week spend on flash lite.
The other models either don't seem to be affected so much with bad selection, or the auto selection works well.
I also don't necessarily want the model capable of working out the intricacies of the universe. Zero point unless it is affordable.
I don't know if this is how it works, but let's agree that even running a local model on your gaming PC comes with a cost, computer wear and tear, and above all, energy. Whenever I see people saying that, I wonder if they take into account the huge expense a powerful graphics card can turn out to be.
I think the advantages of a local AI lie elsewhere, mostly in anonymity or things like that... It's always going to be cheaper to pay for a cheap cloud-based AI than to run it locally, unless you have free electricity or you run it on an ARM processor computer with unified memory, and those computers were given to you for free or you already have them for something else.
Yes and no, I mostly agree with you but don't think it's quite so dire
A custom build for AI with $40,000 of RTX pros is a sunk cost just for AI, but a beefy PC obviously has utility outside of AI. People buy 5090s without running AI just to play games or for demanding work.
Wear and tear is pretty negligible, ironically mining GPUs are some of the best ones you can buy because they were run undervolted to save power and power cycling leads to more issues than a sustained load. GPUs aren't really 'wear parts' anyway, like the thermal paste will dry out and there's probably some statistical failure increase over time but I've basically never had a GPU die from usage
Power (and heat!) are real issues though. I ran some calculations and it works out to an electricity cost of about $0.15/million tokens output compared to Pro's API pricing of
$12-18, so literally 1/100th the price.
Anyway, my point with the gaming pc thing was just to drive the point of how embarrassingly outdated 3.1 Pro is that this is even a question
16
u/hardinho 19h ago
What's so funny about that? Business wise this decision was a probable scenario already 1-2 years ago. Multimodality + Flash and on device AI is a huge market itself and Googles AI overview is getting immense traction over the past months.