WEBVTT

00:00:00.000 --> 00:00:03.945
<v Wren>Older cloud AI models are being systematically removed.

00:00:04.105 --> 00:00:10.350
<v Ash>Right. OpenAI, Anthropic, and Adobe all have 2026 deprecation schedules.

00:00:10.510 --> 00:00:15.680
<v Ash>That includes GPT-4.1, Claude Sonnet 4, and Gemini variants.

00:00:15.840 --> 00:00:19.060
<v Wren>So the push is to newer reasoning models or open-weights.

00:00:19.220 --> 00:00:24.190
<v Ash>Exactly. But for local 48GB hardware, the stack is stable.

00:00:24.350 --> 00:00:31.570
<v Ash>No new models fit that box. gpt-oss-20b and Qwen3-Coder-30B are still the picks.

00:00:31.730 --> 00:00:34.075
<v Wren>What about the big vendor claims this month?

00:00:34.235 --> 00:00:40.655
<v Ash>NVIDIA claims up to 1.9x higher throughput with llama.cpp and vLLM.

00:00:40.815 --> 00:00:49.010
<v Ash>But that's a vendor claim. New TTS models like CosyVoice2-0.5B are only cited secondhand.

00:00:49.170 --> 00:00:50.490
<v Wren>Unverified?

00:00:50.650 --> 00:00:57.670
<v Ash>Secondary or vendor-claim. The tracker's verdict: local tier settled, frontier tier is a price war.

00:00:57.830 --> 00:01:01.300
<v Wren>So the action is migrating off the deprecated cloud models.

00:01:01.460 --> 00:01:06.680
<v Ash>That's the data. See the full, machine-maintained tracker at DreamLab Research.
