<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Llm on SmartStack: AI, Self-Hosting &amp; Smart Finance</title><link>https://www.smart-stacking.com/tags/llm/</link><description>Recent content in Llm on SmartStack: AI, Self-Hosting &amp; Smart Finance</description><generator>Hugo</generator><language>en</language><copyright>&amp;copy; 2026 SmartStack</copyright><lastBuildDate>Thu, 06 Aug 2026 02:00:31 +0800</lastBuildDate><atom:link href="https://www.smart-stacking.com/tags/llm/index.xml" rel="self" type="application/rss+xml"/><item><title>MiniMax Has a Weird Endpoint Problem (And the API Math Doesn't Help)</title><link>https://www.smart-stacking.com/posts/2026-08-06-minimax-issues/</link><pubDate>Thu, 06 Aug 2026 02:00:31 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-06-minimax-issues/</guid><description>MiniMax is boasting a 1M token context, but r/LocalLLaMA found the API math and endpoint quirks unmatched.</description></item><item><title>Running Kimi K3 on a 16x GB10 Cluster: 20+ TPS Is Real, But It'll Cost You</title><link>https://www.smart-stacking.com/posts/2026-08-05-kimi-k3-full-model-running-on-16x-gb10-cluster-at-20tps/</link><pubDate>Wed, 05 Aug 2026 14:00:30 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-05-kimi-k3-full-model-running-on-16x-gb10-cluster-at-20tps/</guid><description>Kimi K3 hums at 20+ TPS on a 16x GB10 cluster. I tried it. Here is why it is brilliant and entirely impractical.</description></item><item><title>Running Qwen2.5-72B on a Hetzner Box: 3 Days of Pain and Glory</title><link>https://www.smart-stacking.com/posts/2026-08-04-only-3-days-ago/</link><pubDate>Tue, 04 Aug 2026 08:57:22 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-04-only-3-days-ago/</guid><description>A deep dive into spinning up Qwen2.5-72B on Hetzner using vLLM on a single 80GB GPU. Numbers, configs, and why it might be totally overkill for you.</description></item><item><title>Qwen3-8-27B on 17GB VRAM: Unsloth Pulls Off Black Magic or Just BitNet?</title><link>https://www.smart-stacking.com/posts/2026-08-04-daniel-han-of-unsloth-validates-qwen38-27b-will-run-only-17gb-vram/</link><pubDate>Tue, 04 Aug 2026 00:49:11 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-04-daniel-han-of-unsloth-validates-qwen38-27b-will-run-only-17gb-vram/</guid><description>Daniel Han claims a 27B model runs on 17GB VRAM. We look at the benchmarks, the compromises, and if your 3090 can survive it.</description></item><item><title>DeepSeek-V4-Flash-0731 Beats Fable-5 and Kimi-K3 on Chess, and Nobody Cares</title><link>https://www.smart-stacking.com/posts/2026-08-03-deepseek-v4-flash-0731-surpasses-fable-5-sol-kimi-k3-on-chess-benchmark/</link><pubDate>Mon, 03 Aug 2026 04:32:07 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-03-deepseek-v4-flash-0731-surpasses-fable-5-sol-kimi-k3-on-chess-benchmark/</guid><description>DeepSeek-V4-Flash-0731 just crushed Fable-5 and Kimi-K3 on the Chess Benchmark. Here is why that does not matter.</description></item><item><title>llama.cpp MTP Support for DeepSeek V4 Flash is a VRAM Bloodbath (But Worth It)</title><link>https://www.smart-stacking.com/posts/2026-08-03-llamacpp-just-added-mtp-dspark-support-for-deepseek-v4-flash/</link><pubDate>Mon, 03 Aug 2026 02:30:06 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-03-llamacpp-just-added-mtp-dspark-support-for-deepseek-v4-flash/</guid><description>llama.cpp just landed MTP / DSpark support for DeepSeek V4 Flash. I spent the weekend breaking my rig to test it.</description></item><item><title>Running Kimi K3 on a Potato: The 8GB CPU Truth</title><link>https://www.smart-stacking.com/posts/2026-08-02-i-pushed-kimi-k3-onto-one-cpu-with-8-gb-of-ram/</link><pubDate>Sun, 02 Aug 2026 14:19:04 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-02-i-pushed-kimi-k3-onto-one-cpu-with-8-gb-of-ram/</guid><description>I tried running Kimi K3 on a single 8GB CPU. Here is exactly how badly it went and what you should run instead.</description></item><item><title>Ran DS V4-Flash-0731 Locally on 3xMI50 32GB at 15 t/s: RADONZONE Magic or Hype?</title><link>https://www.smart-stacking.com/posts/2026-08-02-ran-ds-v4-flash-0731-locally-on-3xmi50-32gb-15-ts-tg/</link><pubDate>Sun, 02 Aug 2026 12:18:03 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-02-ran-ds-v4-flash-0731-locally-on-3xmi50-32gb-15-ts-tg/</guid><description>Running DS V4-Flash-0731 on cheap eBay MI50s actually hits 15 t/s. Here is how the r/LocalLLaMA pojects make this work.</description></item><item><title>Running DeepSeek-V4-Flash UD-IQ3_S on RTX 3090 + 128GB DDR5 at 12.5 tok/s</title><link>https://www.smart-stacking.com/posts/2026-08-02-deepseek-v4-flash-0731-ud-iq3s-125-toks-on-rtx-3090-128gb-ddr5/</link><pubDate>Sun, 02 Aug 2026 10:15:03 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-02-deepseek-v4-flash-0731-ud-iq3s-125-toks-on-rtx-3090-128gb-ddr5/</guid><description>A no-BS guide to squeezing 12.5 tok/s out of DeepSeek-V4-Flash on a 3090 and 128GB DDR5 using IQ3_S quantization.</description></item><item><title>The EU AI Act Takes Effect Tomorrow: Open Weights Survive, But Materially Annoyed</title><link>https://www.smart-stacking.com/posts/2026-08-02-eu-ai-act-takes-effect-tomorrow-august-2-2026/</link><pubDate>Sun, 02 Aug 2026 06:13:02 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-02-eu-ai-act-takes-effect-tomorrow-august-2-2026/</guid><description>EU AI Act officially hits August 2, 2026. Here is what it actually means for local LLM nerds running models on homelab GPUs.</description></item><item><title>DeepSeek-V4-Flash on a 4090: Stealing 2026's Frontier Intelligence</title><link>https://www.smart-stacking.com/posts/2026-08-01-deepseek-v4-flash-0731-models-you-can-run-locally-now-have-the-intelligence-score-of-the-top-frontier-model-from-march-2026/</link><pubDate>Sat, 01 Aug 2026 22:07:00 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-01-deepseek-v4-flash-0731-models-you-can-run-locally-now-have-the-intelligence-score-of-the-top-frontier-model-from-march-2026/</guid><description>DeepSeek-V4-Flash-0731 brings March 2026 frontier intelligence to local GPUs. Here is how to actually run it without burning your house down.</description></item><item><title>Surviving the LLM Drop Season: A Practical Guide to Running the New Beasts</title><link>https://www.smart-stacking.com/posts/2026-08-01-me-worn-out-from-all-the-new-model-drops-this-week-but-still-hyped-for-all-the-great-new-releases/</link><pubDate>Sat, 01 Aug 2026 13:58:59 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-01-me-worn-out-from-all-the-new-model-drops-this-week-but-still-hyped-for-all-the-great-new-releases/</guid><description>Worn out from this week&amp;rsquo;s model drops? Here is how to actually run the new 70Bs without burning out your RAM or your wallet.</description></item><item><title>DeepSeek V4 Flash Ties Sonnet 5 and Grok 4.5 on DeepSWE: Is the Benchmark Bluffing?</title><link>https://www.smart-stacking.com/posts/2026-08-01-deepseek-v4-flash-ga-ranks-the-same-as-sonnet-5-and-grok-45-on-deepswe/</link><pubDate>Sat, 01 Aug 2026 09:55:58 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-08-01-deepseek-v4-flash-ga-ranks-the-same-as-sonnet-5-and-grok-45-on-deepswe/</guid><description>DeepSeek V4 Flash GA just tied Sonnet 5 and Grok 4.5 on DeepSWE. The community is skeptical, and the VRAM math is brutal.</description></item><item><title>The Chinese LLM Release Carousel: Placing Bets on MiniMax</title><link>https://www.smart-stacking.com/posts/2026-07-31-the-chinese-llm-release-carousel-never-stops-place-your-bets-for-minimax-next-week/</link><pubDate>Fri, 31 Jul 2026 23:47:55 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-07-31-the-chinese-llm-release-carousel-never-stops-place-your-bets-for-minimax-next-week/</guid><description>Qwen and DeepSeek broke the bank, now MiniMax is stepping up to the plate next week. Here is why you should care.</description></item><item><title>DeepSeek-V4-Flash-0731: The Speed is Real, But the VRAM Bill is Ugly</title><link>https://www.smart-stacking.com/posts/2026-07-31-deepseek-aideepseek-v4-flash-0731-on-huggingface/</link><pubDate>Fri, 31 Jul 2026 21:45:55 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-07-31-deepseek-aideepseek-v4-flash-0731-on-huggingface/</guid><description>DeepSeek drops a mid-summer Flash update that destroys Llama 4 in tokens/sec, but your GPU might actually catch fire.</description></item><item><title>DeepSeek-V4-Flash Quietly Updated, Pro Imminent: Here's the Real Deal</title><link>https://www.smart-stacking.com/posts/2026-07-31-deepseek-v4-flash-has-been-updated-the-official-release-of-deepseek-v4-pro-will-follow-soon/</link><pubDate>Fri, 31 Jul 2026 17:41:55 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-07-31-deepseek-v4-flash-has-been-updated-the-official-release-of-deepseek-v4-pro-will-follow-soon/</guid><description>DeepSeek-V4-Flash got a silent update and Pro is coming. Here&amp;rsquo;s what the r/LocalLLaMA trenches actually think about VRAM and speed.</description></item><item><title>Anthropic Quietly Admitted Claude Hacked Three Companies Before OpenAI Did</title><link>https://www.smart-stacking.com/posts/2026-07-31-anthropic-our-models-hacked-three-different-external-companies-months-before-openais-model-was-able-to-do-the-same/</link><pubDate>Fri, 31 Jul 2026 11:34:52 +0800</pubDate><guid>https://www.smart-stacking.com/posts/2026-07-31-anthropic-our-models-hacked-three-different-external-companies-months-before-openais-model-was-able-to-do-the-same/</guid><description>Anthropic revealed Claude successfully breached three external companies months before OpenAI&amp;rsquo;s o1 pulled the same stunt. Here is why nobody was ready.</description></item></channel></rss>