Hacker Newsnew | past | comments | ask | show | jobs | submit | more aray07's commentslogin

effort level is separate from tokenization. Tokenization impacts you the same regardless.

I find 5 thinking levels to be super confusing - I dont really get why they went from 3 -> 5


i think the new qwen models are supposed to be good based on some the articles that i read


anthropic’s pricing is all based on token usage

https://platform.claude.com/docs/en/about-claude/pricing

So if you are generating more tokens, you are eating up your usage faster


are you okay with paying more for your services without any perceived improvement in the service itself?


That's been a constant for my entire adult life.


yeah thats is my biggest issue - im okay with paying 20-30% more but what is the ROI? i dont see an equivalent improvement in performance. Anthropic hasnt published any data around what these improvements are - just some vague “better instruction following"


Its enshittificating real fast. They'll just keep releasing model after model, more expensive than the last, marginal gains, but touted as "the next thing". Evangelists will say that they're afraid, it's the future, in 6 months it's all over. Anthropic will keep astroturfing on Reddit. CEOs will make even more outlandish claims.

You raised a good point, what's a good metric for LLM performance? There's surely all the benchmarks out there, but aren't they one and done? Usually at release? What keeps checking the performance of those models. At this point it's just by feel. People say models have been dumbed down, and that's it.

I think the actual future is open source models. Problem is, they don't have the huge marketing budget Anthropic or OpenAI does.


This is most likely trajectory I fear. It reminds me a lot of Oracle, where they rebrand and reskin products just to change pricing/marketing without adding anything.


Win 10, win 11, all the recent macOS,… could have been released as features and not new products


My only problem lumping them in is a huge part of those products are the UI and UX. It’s what you interact with as a user. So, a reskin with no new features is a valid version in my opinion.

It’s not always for the best mind you, but I do think major changes to the GUI or UX do deserve a version indicator given they are relatively infrequent. Without some requirement of course.

Nobody buys Oracle because of the UI or upgrades because of a new UI. It’s the underlying tool that matters and if my ugly version is working there is no incentive to upgrade to a prettier version.


The other thing is most people don't really care about price per token or whatever but how much it will cost to execute (successfully) a task they want.

It doesn't matter if a model is e.g. 30% cheaper to use than another (token-wise) but I need to burn 2x more tokens to get the same acceptable result.


isn’t caveman a joke? why would you use it for real work?


yeah opus 4.7 feels a lot more verbose - i think they changed the system prompt and removed instructions to be terse in its responses


yeah similar for me - it uses a bunch more tokens and I haven’t been able to tell the ROI in terms of better instruction following

it seems to hallucinate a bit more (anecdotal)


I had it hallucinate a tool that didn't exist, it was very frustrating!


Anthropic intruduces fake tool calls to prevent distillation of their models. Others still distill. Anthropic distils third party models. Claude now hallucinates tools.

Brilliant.


yeah i am still not clear why there are 5 effort modes now on top of more expensive tokenization


choice is often a great dark-pattern (lack of choice is too but...). Choices generally grow cost to discover optimality in an np way. This means if the entity giving choice has more ability to compute the value prop than the entity deciding the choice you can easily create an exploitive system. Just create a bunch of choices, some actually do save money with enough thought but most don't, and you will gain:

People that think they got what they wanted, the feature is there!, so they can't complain but...

People that end up essentially randomly picking so the average value of the choices made by customers is suboptimal.


Once you've seen a few results of an LLM given too much sway over product decisions, 5 effort modes expressed as various english adjectives is pretty much par for the course


good point - i analyzed text tokenization. will run some experiments to see how visual tokenization has changed


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: