What good is fable if it hits the 20x limit in one prompt before completing the task?
Login to reply
Replies (8)
Its why you should not dismiss open source models. Maybe its not fable tier, but they are much cheaper.
It's all about getting you to pay for tokens.
I think this probably has something to do with the way Anthropic does cache reads and writes.
Insane that Fable doesn't write to cache after the initial system prompts. Notice that it wrote 4.9k tokens to cache when the session began and never again. I feel rugged! This planning session cost me 8.9k sats. 😭
Deepseek V4 Pro caches everything and costs much less, proportionally. (see image below)
ps: this is routstrd top that shows cache r/w data

View quoted note →

Buy dips not hype
which one do you recommend?
Depends on your use case and if you wish to run something locally or if you want an online api.
If its the latter Kimi is probably one of the best.
For local its between Gemma4 and Qwen3.5/Qwen3.6 .
Qwen3.5 also has larger versions.
Because it can hold both a thesis and an antithesis without melting under basic neurotypical thinking. Fable is a beast for researching complex topics