The "DeepSeek V4 Flash + GLM-5.2" Combination Is Currently Enough for Me (and Will Probably Be Enough for You Too)

(Antonio V. Franco) When I turn on my computer in the morning, one of the first things I do is open whichever Ollama Cloud account I’m currently using and check my consumption. I have three Ollama Cloud accounts, each on the US$20 plan, which gives me considerably more peace of mind regarding usage, even though I still can’t afford to consume resources recklessly. ...

July 2026 · 5 min · Antonio V. Franco

GLM-5.2: I thought Sonnet 4.5 and similar open models were enough, but…

(Antonio V. Franco) Honestly, it had been a long time since a model surprised me like this — not since the end of 2025. No, I do not mean it surprised me because I liked a response or was impressed by a benchmark. No, and in this article you will understand where I am going with this. ...

June 2026 · 5 min · Antonio V. Franco

Agentic AI, SLMs, and Why Models Above US$0.50 Output per 1M Tokens Are Equivalent to Burning Money

(Antonio V. Franco) June 2026. If you are still paying more than fifty cents per million output tokens in any agentic pipeline, forgive my frankness: you are literally burning money. No, this is not an exaggeration on my part, but simply the relationship between API costs and absurd usage — because, indeed, LLMs should be used absurdly, whether you are a solo researcher like me or a startup. And here, the reason has a name, surname, and version: DeepSeek V4 Flash. And what it represents goes far beyond excellent benchmark results — it represents the consolidation of a trend that demonetizes the already psychologically consolidated commercial models and pushes the center of gravity toward SLMs, the highly specialized Small Language Models. ...

June 2026 · 5 min · Antonio V. Franco