46 comments

  • giancarlostoro 1 hour ago
    Call me crazy but:

    VRAM & Memory Requirements by Precision

    • FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).

    • INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).

    • INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)

    VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.

    Even so why would anyone not sleep on a model they cannot run?

    • kristopolous 1 hour ago
      Seriously, if a single politician stepped forward and said "i'll bring down ram prices" they could then shoot a puppy and call me a slur and I'd still go out and doorknock for them.

      Memory companies have price fixed multiple times. They've paid hundreds of millions in fines. wikipedia even has a page on it. https://en.wikipedia.org/wiki/DRAM_industry_price_fixing.

      Look at the financials of these companies, they're all making obscene margins and do they plan to increase production? No. Micron is doing a stock buy back to pump the price of their share.

      The Micron CEO just recently said this is the exact plan https://www.theregister.com/systems/2026/10/01/ram-supply-se...

      There's sanctions, tarrifs, and a DOJ who doesn't give a shit. Until we can fix that the insanity will continue. Phones will be unaffordable. Laptops will be obscene. Gaming consoles will be thousands of dollars. Desktops will be dead.

      If you're waiting for some David Ricardo equation to happen, tough cookies, it's not coming.

      The market is legally locked down and we're in hostage pricing mode.

      And what's the story? You can't afford electronics because we're using it to build robots to take your job? I mean ...

      Nobody is coming to save us. That's our job.

      • phil21 58 minutes ago
        > do they plan to increase production? No.

        Micron has 3 brand new fabs currently under construction, 2 Boise, 1 in New York as the first of 4 planned for a campus.

        Plus expanding other existing facilities.

        These things take ~3-5 years from breaking ground to full production. You'd have had to anticipate the current demand years before it happened in order to be bringing production on-line before 2030 or so.

        Samsung and HK Hynix also have fabs under construction and planned.

        CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.

        Not much you can really do to wish for more fabrication to exist on any timeline not measured in fractional decades.

        Could they do more and react quicker? Probably, but everything I've read on the subject seems to point to 3 years is absolute bare minimum if you happen to have a shovel ready project with the land bought, local permitting completed, infrastructure extended to the site, and a skilled workforce already in place. They could suspend buy-backs/dividends today and dump it all into building production and there would be no material impact until around 2030.

        > The Micron CEO just recently said this is the exact plan

        CEO simply stated the demand pressure will not go away through 2027, and supply will not increase until around 2028 when currently under construction fabs start shipping volume. The article does not support your statement.

        • ttul 31 minutes ago
          Stanford tracks RAM prices in this nice little site: https://dam.stanford.edu/memory-prices.html

          Costs did go nuts, but there are signs of easing in the market of late. CXMT is starting to have an impact and priced will probably fall in 2027.

        • cogman10 15 minutes ago
          The second Micron boise fab hasn't even broken ground yet, they are still working on the first one. So don't expect these things to be completed in parallel.

          Some of my family is pretty happy, though, with the job security as they are pretty convinced these projects are all going to take much longer than what's being stated publicly. Micron is saying the first chip from the new fab will be in 2027... though they also predicted it'd be 2026. The date seems pretty slippy.

          • dboreham 10 minutes ago
            Anyone who has been around the semiconductor industry since the last century will remember various huge fabs e.g. in Arizona that were partially built but never finished due to oversupply by the time the walls and roof were done.
        • BizarroLand 31 minutes ago
          Yeah, but why would they make consumer memory when HBM for GPUs is much more profitable?
          • kristopolous 14 minutes ago
            Capitalism eats itself this way. Second and third order effects will collapse the demand.

            You need to keep the market healthy, not some insane Bitcoin style HODL pump - that's how you get wrecked.

            I mean I'm not a neoclassicalist but I've read all of them. I'm in consensus with them here. There's a bunch of theories on what a healthy market is but what we're currently seeing matches none of them.

            It's short term profitable but long term disastrous, especially in a world where new mathematics and techniques could literally collapse the demand overnight.

            Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!

            Some clever trick about how attention heads and context Windows work could potentially slash a bunch of requirements by giant margins and all they're doing is firing the starting gun at that global race with every obscenely priced unit they sell.

            But if prices were reasonable, this wouldn't be an apocalypse. It'd be fine. Consumers wouldn't rush to 64GB, they'd say " Cool I can multitask now at 256" or " great I can do horizontal scalability' or something else.

            But no they created the market conditions so now what would happen is the consumer will immediately flip the 192GB they don't need on eBay, hoping to snatch a profit before the prices tank and the second hand market will be flooded the rug will be pulled out from the luxury pricing and everyone will get screwed.

            This has happened in electronics markets before. Many times.

            When Engels talked about the grave diggers of capitalism they were looking at it through a 19th century labor/manufacturing lens but arguably this same dynamic is at play here.

            • Analemma_ 9 minutes ago
              What "second- and third-order effects" do you suppose will collapse the demand for RAM? The people complaining most loudly about RAM costs are the people who want to run local models; if that becomes popular it will supercharge RAM demand, because locally-hosted models can't parallelize runs from many users the way cloud-hosted ones can. I don't see any slackening in RAM demand at any point in the foreseeable future, even if the big AI companies all go bust.
      • bob1029 0 minutes ago
        [delayed]
      • ashdksnndck 51 minutes ago
        RAM manufacturers are bidding against NVIDIA and everyone else for the same constrained supply of EUV machines. And it takes years to build more fabs. Micron has multiple fabs coming online in 2027 and 2028.
      • m463 1 hour ago
        > "i'll bring down ram prices"

        wonder what voting would be like?

        gamer vote ++

        datacenter hater vote --

        datacenter lobby ++

        micron lobby --

        • rezonant 7 minutes ago
          Yep, that's all the voting blocs.
        • mwambua 57 minutes ago
          Wouldn’t cheaper memory make it easier to bring compute out of data centers and onto consumer hardware?
          • TeMPOraL 38 minutes ago
            Datacenter haters will read this as "that's still evil AI", and everyone else hopefully can count and understands it'll be worse for environment.
      • neya 13 minutes ago
        > they could then shoot a puppy and call me a slur

        I know it's just a figure of speech, but damn. I laughed out aloud in public just reading this.

      • xyzsparetimexyz 56 minutes ago
        Neither political party cares at all about memory pieces get real lol
        • kristopolous 48 minutes ago
          wait until holiday shopping...it affects the price of almost everything with a battery or power cord.
      • gchamonlive 1 hour ago
        [flagged]
        • Analemma_ 1 hour ago
          I don't think RAM vendors have formed a cartel and I think this is knee-jerk anger without any thought. RAM is a commodity product with massive upfront capex costs, and those always have boom-and-bust cycles. At various points in the 2010s and 2020s RAM vendors were getting eaten alive by a supply glut, this would not have happened if they were a cartel.

          Is it really so hard to believe that RAM prices are up because demand is simply exceeding supply, especially in a market where additional supply takes years and billions of dollars to come online? There's no need to posit cartel behavior and a fair amount of evidence that there is none.

          • kristopolous 35 minutes ago
            The AI boom started in 2022. Prices rose THREE years later after 2025 sanction and tariff style legislation to protect the market during a price hike.

            I got a 4090 in 2023 for 1600, a 5090 in 2025 for 2000 with 256 DDR5 for about $1,000 ... and then, after some protectionist legislation passed, these prices quickly shot to the moon.

            Connect the dots.

            • Analemma_ 2 minutes ago
              Man I think you're just spewing word salad and a lot of what you've written is either wrong or not even wrong. The AI hype really got started in 2022, but hype on social media doesn't mean anything for RAM prices, only real buildouts do that. They rose pretty steadily until OpenAI revealed their shenanigans re: locking up a ton of supply from two different vendors with secret contracts, and that's when the takeoff really started. This is definitely scummy behavior from OpenAI (big surprise), and I actually think they arguably should see an antitrust investigation for that (not that that will ever happen), but OpenAI is a buyer; that's not the same thing as the vendors forming a cartel.

              You can't say "connect the dots" at the end of a raving, mostly-incorrect post and act like you've made an ironclad argument.

    • jauer 3 minutes ago
      This “blame sama for memory prices” meme is so tired.

      He gave demand signal so many times years ago and was mocked for it and now we have the consequences of industry not taking him seriously.

    • petu 1 hour ago
      There's no BF16, original full quality weights are quantized already and 510GB.

      Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.

      Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).

    • ManuelKiessling 9 minutes ago
      Thanks for the data!

      Allow a question from someone who’s only got a very vague idea of how this kind of stuff works behind the scenes: say I rent usage of this model through one of the many LLM hosting providers out there, and let‘s assume I use it extensively through something like Pi or OpenCode and vibe code away all the time, keeping the hosted model occupied as much as I can, happily burning my credits.

      Does that mean that there is a hardware cluster as described by you above that is crunching away just for me?

      So at FP16, I alone keep a 1,664 GiB system occupied all the time?

      • DrammBA 0 minutes ago
        No, a cluster can server multiple users at the same time, providers cap the tok/s so that one cluster can run inference on multiple inputs at the same time. OpenAI with their new ultrafast mode is probably reserving the whole cluster or prioritizing requests of ultrafast users above others with a higher tok/s hence the high price and high speed.
    • crossroadsguy 21 minutes ago
      I did somet math and completely gave up on the idea of trying any worthwhile local model and figured I'd rather pay the 15-30 USD per month via subscription and/or API key combos for years than buying a local setup which might go out of date very fast, if it doesn't goes kaput just out of warranty. I won't be surprised if RAM scarcity is an concerted effort to herd people towards the remote models :)
    • keammo1 32 minutes ago
      The article isn't just about running locally though. The author is saying it's super cheap to run the model through Opencode Go (and presumably OpenRouter etc.) Personally I'm always most excited by models I can actually run locally, but even these huge open source models open up the competitive landscape for companies to let you call models via an API or just lease compute. And they don't have to charge you to offset research, training, huge staffs of the best minds in the world, crazy PR etc. I think that's a big win for customers and buts competitive pressure on the frontier labs as well.
    • apitman 36 minutes ago
      > Even so why would anyone not sleep on a model they cannot run?

      Because it's an open model so providers compete on price.

    • girvo 44 minutes ago
      Not quite: not all of this needs to be in VRAM

      It has a set of n-gram tables which you can stream from system RAM or even NVMe

      That said it’s still quite big! I can’t fit it on my DGX Spark, though I believe you can if you have two?

      • rsolva 18 minutes ago
        I have access to two and will explore this the coming weeks.
    • ByteAtATime 1 hour ago
      I think, considering the size of this model, it's closer to a Pro than a Flash on everything other than speed
    • anvuong 57 minutes ago
      I just un-retire my pair of 1080Ti for some small models development because the current GPU prices literally make me sad.
    • cookiengineer 38 minutes ago
      It's dangerous to go alone. Take this: [1]

      I reimplemented most of the features of the Deepseek v4.1 flash paper (apart from quantization aware training which doesn't make sense because my implementation uses float32 precision anyways)

      I'm currently learning how to distill reasoning traces (check my other github repositories) but I think that a locally selfhostable deepseek is possible with my mixture of experts sharding mechanism. I decided to optimize everything for CPU parallelization, with the idea that the KV cache and meta model have to run from CPU RAM anyways, so the experts can also be loaded/unloaded at runtime if needbe, to save more RAM.

      My assumption is that the KV cache optimizations in combination with the CED and compressed attention features are the reason why v4.1 flash has so few hallucination problems and such a strong self-lookup/thinking behavior. But that's more a gut feeling, need to evaluate and test this more thoroughly.

      Anyways, would love to see someone train this on their own datasets. Currently my pipeline is kinda optimized for parquet and zim files.

      [1] https://github.com/cookiengineer/gonano

    • CamperBob2 29 minutes ago
      You can run it locally for the price of a decent car, or run it on somebody else's hardware at vast.ai or a similar provider for much less. What's not to like?

      No, you won't get frontier-level intelligence on a 1070Ti. Yes, it should be illegal to do what Altman did. Since we clearly don't live in the best of all possible worlds, we need to settle, and DS4.1 Flash is a good place to do that.

      For tasks that don't require vision to support, I personally like the NVFP4 quant of GLM 5.3 from Local Inference Lab better, but they are both well beyond awesome.

    • nullc 54 minutes ago
      The bulk of its weights are natively MXFP4. And engram values don't need to be in vram.
    • functionmouse 1 hour ago
      one can make a fine gaming pc for ~$350

      1660 ti, 4790k, 16gb ddr3

    • holoduke 1 hour ago
      He doesn't ruin the cost of memory. Advances in memory size and speed are now in full speed mode. Expect drastic increase in the upcoming years. Big factories are in the making and planned. Gigalab in the US and many others in the east. Since 2010 we have computers with 16gb as being normal. Finally we are moving into a new era where the standard will be 64gb next year and 128 in 2028. Hopefully we reach 1tb in 2030.
      • CorrectHorseBat 1 hour ago
        I've read the exact opposite, vendors are reducing the standard from 16GB back to 8GB
        • holoduke 28 minutes ago
          That's only temporary till production meets demand again.
          • CorrectHorseBat 14 minutes ago
            Which is not going to happen in the upcoming years
    • liuliu 1 hour ago
      What are you talking about? The model is native NVFP4, why you run it at any precision higher than that?
  • mlinsey 50 minutes ago
    I'm paying for the heavily-discounted subscriptions, not the API rates. There isn't really a cost gap for me. DeepSeek doesn't have a subscription to compare to, but when I compared the GLM 5.3 usage I got from a $100/mo Z.ai subscription compared to Opus 5.5 on a $100/mo Claude subscription, there wasn't a big gap. And GLM 5.3 is very clearly not a frontier model (deepseek v4 seemed a lot closer, but I didn't use it enough to really say for my workloads).

    I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).

    Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.

    • apitman 13 minutes ago
      The tightening of subscription value has already begun. dsv4.1f is already worth paying for at market API prices. Maybe it goes to 2x because apparently no one has figured out how to match DeepSeek's insane caching efficiency, but I don't see it getting much worse than that.

      Plus you can also get dsv4.1f subsidized. OpenCode Go gives 4x if I understand their pricing correctly. Anecdotally, I feel like I get way more out of my $10/mo OpenCode Go sub for the price than my $20/mo ChatGPT, even using gpt-6.1-sol high which is very cheap, and I have yet to convince myself dsv4.1f is a worse model.

      • rapind 1 minute ago
        It's really not. It's getting closer, and it's a great model, but it is not more value per task than the frontier subscriptions. Don't be swayed by the token costs, it's very chatty, like 3x more tokens for the same task as sol. I used dsf 4.1 full time for about a week.

        It blows frontier API pricing out of the water, but again, look at cost per task, not token usage. Still easily wins though for my work.

        I do think it's the most viable alternative I've seen so far, and that applies pressure to the frontier models. Should subscription prices hike or become unavailable for some reason, I know what I'll be using.

        When pricing this, it's important to consider whether or not you want to opt out of data training. You won't get the advertised rate.

  • apitman 5 minutes ago
    > With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited

    My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).

    I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.

    This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.

  • arush15june 2 minutes ago
    I am 4.1 maxxing on commandcode GOAT Plan + api rates with oh my pi for the last 4 weeks, it's absolutely amazing and crazy fast, it's alright if it makes a mistake, I have enough time to iterate again, I have also added an advisor layer of mimo 2.6 pro which does make it a notch smarter. Getting haiku 5.5/sonnet5.5 to work on plans and letting 4.1 flash work through it is helping a ton too.

    I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.

    Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.

    And it never says no for cyber tasks so that's a big win

  • lmf4lol 1 hour ago
    Oh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.

    Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.

    But as a main driver. I love flash. And it brought our bill down by A LOT :D

    • crossroadsguy 15 minutes ago
      What is the cost of access like for DeepSeek-v4.1-flash, compared to GLM-5.3-flash via ZAI's Coding Plan? Because that's what I use; and often hit the "wait". I wouldn't mind trying a new model subscription or even API access which hits around glm-5.3-flash level weight class (which seem to be enough for me; with quite some human suprvision and nudging) but gives muuuuuuch moooore tokens for the same price.

      (And what are the preferred providers?)

    • PcChip 27 minutes ago
      >We run all our Personal Assistants now on flash

      are you worried about sending all your data to third parties, especially if they're in different countries?

      • techmunky 7 minutes ago
        Not shilling for them but Ollama cloud hosts domestically with ZDR afaik. I run 95% of my open weight inference through them. The rest goes through Opencode Go $10 plan (which is enough to run 3 hermes agents on DSF 4.1 and leave plenty of left to experiment with when new models drop).
        • octoberfranklin 2 minutes ago
          Just a reminder that any API using a Cloudflare TLS certificate isn't ZDR.

          The model engine provider might be ZDR, but the service as a whole isn't.

      • crossroadsguy 14 minutes ago
        I have asked OP that question but I think there are providers who are not in China and they just host the model/inference.
      • yieldcrv 13 minutes ago
        Just use a provider hosting it in your country especially if your country has major data centers then its the same as using Anthropic or GPT of GCP Model Garden or AWS Bedrock

        nobody here is talking about running frontier level intelligence locally so if you’re Chinaphobic and prefer layers of corporations siphoning your data in between you and the party there are plenty of options instead of directly to the party

    • aftbit 57 minutes ago
      Have you compared it against actual SOTA models like latest Fable or Astra?
      • sneurlax 50 minutes ago
        Of course there's still a huge performance gap

        but DS 4.1 Flash is good enough for most tasks

  • p1necone 37 minutes ago
    I have a pretty large, complex project I've been building with heavy AI use (new language + compiler). I was following a 'strong model as orchestrator launching cheap models as implementers' pattern, but I recently trialled just using Deepseek-V4.1-Flash as the model for both layers because of the cost savings (with mimo v2.6 flash on code review agents for some decorrelation).

    I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.

    However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.

    • rspeele 28 minutes ago
      On the Claude side of things I was previously following "strong model directs weak" with Fable directing Opus/Sonnet (its choice per-task). Since Opus 5.5 came out I've just been having Opus direct Opus.

      The sub-agent separation is still valuable to keep context clean for the orchestrator, but I just have no reason to use Sonnet as the grunt-work implementer because I'm finding it hard to run out of tokens with Opus 5.5 on a $200 subscription plan. It's really really good at subjective quality of work per token used.

      • p1necone 26 minutes ago
        I would probably go that route if I could use other harnesses with claude models, but I don't want to be locked in to claude code, and their API pricing (which you need to use it with other harnesses) is so much higher than subscription.
  • hmontazeri 1 hour ago
    I had the same experience using ds 4.1 last couple of weeks. It’s insanely good for the price. I’m doing mostly web dev it excels at everything I throw at it. The pricing is ridiculous. I canceled my gpt subscription and haven’t looked back hope the pricing stays like that. I almost never need a better model. I still keep my Claude 20$ sub for now but I feel like one more iteration and I won’t need even that anymore I hardly use it
    • jacquesm 1 hour ago
      If DS4.1 impresses you I would be really interested to see your comparison to GLM 5.3. I switched from the one to the other and even if GLM 5.3 is a bit slower I don't think I'll be going back.
      • ctolsen 1 hour ago
        GLM 5.3 is very impressive and definitely better, but it also at least 4x the price.

        On that note I’ve been subbing in MiMo-2.6-pro when cost is an issue, which is super cheap and also performing really well.

    • pdhborges 53 minutes ago
      What inference provider are you using?
  • zug_zug 20 minutes ago
    I did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.

    That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.

  • alex-moon 48 minutes ago
    I think because we're all just using it thinking we have found the "model for me" and never mentioning it to anyone because what would we say? It's good. It's a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.
  • RGS1811 51 minutes ago
    This model finally got me off my Claude Max subscription. I’ve found it superior to Opus 5.5 in certain use cases, and certainly faster.

    I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.

  • swiftcoder 1 hour ago
    I think the interesting provider to cross-check this assumption with here is Meta, who is clearly freaking out, and is currently providing Muse 1.3 even cheaper so long as you are willing to share data with them
    • 9dev 16 minutes ago
      Whatever the question, Meta is the wrong answer.
  • jbellis 25 minutes ago
    I built mjolnir in large part so I could have Opus manage DeepSeek Flash subagents. It's phenomenal and extremely light on the Claude tokens. https://github.com/BrokkAi/mjolnir/

    And yes, Opus is enough smarter than DSF that it's worth the extra steps. This ranking is from live tickets, no contamination: https://slopcop.com/power-ranking

  • wg0 1 hour ago
    While using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.

    I realized that mistake and guided DeepSeek where it should be.

    Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.

    • sampullman 1 hour ago
      Do you mean Fable 5.1? Or Opus 5.5? I'm not sure what you're working on but for me DS 4.1 flash isn't nearly at their level. For the price it's obvious very impressive, though Luna 6.0 is excellent too.
      • hirako2000 55 minutes ago
        The problem with benchmarks and proprietary models is that one day a model is best at doing X, another day that's not so sure. And anyway, we are throwing the same X.

        I've found supposedly smaller and, less performant models do better on certain tasks. I end up using several models, sticking to what my unconscious statistical observations tell me to use for the kind of task at end.

      • wg0 45 minutes ago
        Fabble 5.1.
  • aguilaair 1 hour ago
    What about MiMo v2.6 Pro? It’s throughput is slower by default (UltraSpeed is faster than DS4.1F) but is above the pareto line, and cheaper.

    see https://artificialanalysis.ai/models/releases/comparisons?co...

    • patresh 37 minutes ago
      Technically yes, but has been reported to be quite benchmaxxed. In practice Deepseek Flash 4.1 and GLM 5.3 might therefore still outperform Mimo 2.6 pro.
  • simpaticoder 1 hour ago
    The question seems rhetorical but I think there are two reasons in some combination. First is it there is some awareness lag here. That lag can be on the producer and consumer side. Software enterprises are pretty slow to adopt new things and slow to try new things so they might only be aware of openai and Claude as options. Plus there are some scariness because deep seek is a Chinese model and therefore export restricted - never mind that there are American in European providers.

    The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.

    • agoodusername63 1 hour ago
      I think it also has a bit to do with the AI sector of tech still moving at lightning speed.

      Theres already models that outdo DS 4.1 flash in cost/performance. Luna 6 on max effort for example. Luna also doesn't care what time of the day it is for cost calculation.

      And I'm sure by the time people ask why Luna 6 is being slept on there will be another cost/performance king

  • f6v 11 minutes ago
    My anecdotal experience is that I can’t even trust DS4Pro let alone Flash. I always have to have Sol reviewing the code.
  • elmer2 54 minutes ago
    DeepSeek isn't even on my mind. I use the frontier models and can get the best in the industry for a relatively cheap price.
    • qwerpy 17 minutes ago
      Yeah. $100 for Claude just about gives me all the usage I want, as a more or less full-time hobbyist having it work in the background most of the day. I was trying to economize by having a local LLM, then Deepseek, then Cursor/Grok, and then I got a taste of Opus 5.5 and I simply cannot go back to having to carefully spec things out and double-check work. I just let it decide, Opus or Sonnet for the next task, and I get almost perfect results. Probably similar with OpenAI's models.

      The token-equivalent monthly spend is > $5K+. If Deepseek's token cost is 20x cheaper, that's $250/mo, and I'd be spending a lot more of my brainpower babysitting it and getting worse results.

      For business/team accounts that pay per-token, maybe I can see the "freaking out" being warranted on the part of the fronter labs. But as long as they're willing to subsidize their end-user subscriptions, I'm not going to move off of them until the alternatives are truly at their level.

  • wildster 28 minutes ago
    I like GLM 5.3 Flash, it seems good enough for coding features if you have a good structure and a good AGENTS.md
    • david-gpu 24 minutes ago
      Don't you run into it sometimes outputting a few Chinese characters, or Cyrillic, for no apparent reason? I fear it writing some nonsense in the code or the terminal. DeepSeek V4.1 Flash doesn't seem to do that.
      • HeavenFox 19 minutes ago
        To be fair even OpenAI's and, to a lesser extent, Anthropic's models do that sometimes
  • aszen 31 minutes ago
    Because subscription plans are cheaper, only enterprises paying per tok pricing should be freaking out
  • browningstreet 1 hour ago
    What would freaking out look like, or is this just a stupid bloggish title flourish?

    Is OpenAI coming in $20B under a sign of "freaking out"?

    • jerf 1 hour ago
      It would look like major chaos in the markets.

      People tend to conflate the question "is AI a useful technology?" with "are the AI companies going to do well?" but they're surprisingly separated in practice, with either one able to be true while the other is false. There is a lot of money tied up in a lot of hardware with a lot of loans made against that hardware as collateral all based on the assumption that AIs are going to need more and more and more and more hardware and whoever has the hardware wins. If a much better model comes out that requires vastly less hardware, or even more accurately, merely charges vastly less than the current AI companies, then to a first approximation (barring Jevon's paradox, and bearing in mind there's no timeline guarantee on that) all that hardware becomes much less valuable for being grotesquely oversupplied relative to what is necessary, and even though that would generally make AI objectively more useful than it was before, it would cause mass financial chaos in the markets.

      The markets need a very particular rate of progress. It isn't entirely clear to me that it's even a possible rate of progress, it may be overconstrained, but they certainly don't have plans for the AI models to get commoditized on the timeframes of these vast, vast array of loans being made against hardware as collateral. Spend a metric shit ton of money to kill all your competition then charge monopoly rent on the one thing absolutely everyone needs doesn't work if you can't economically "kill all your competition" because the economics favor them in the spending spree.

      And then, based on the fact that this is not even remotely complicated logic, there are plenty of people who are fully aware that they have a lot of money tied up in not running around telling everyone how wonderful the cheap models have become.

      • hirako2000 50 minutes ago
        It's also unclear whether those who approved those loans understand GPUs depreciation. In any case, progress in software but also hardware could bring chaos and ruin their house of cards.
        • pessimizer 39 minutes ago
          > It's also unclear whether those who approved those loans understand GPUs depreciation.

          This also assumes heavy utilization, though. If there's heavy utilization, it might mean they're doing well. If they're all spinning, it's time to raise prices.

          • hirako2000 5 minutes ago
            Only if utilization isn't at a loss. Right?
    • efficax 1 hour ago
      they should be freaking out because every time the chinese labs or non "frontier" labs release a model that is only a few months behind and much cheaper than the openai/anthropic models it shows that they don't deserve their valuations
      • WJW 11 minutes ago
        Perhaps, OR it might be that most people in the markets (think that they) are not all that exposed to the valuation AI labs and so their eventual collapse doesn't matter.

        Or perhaps they consider the upside from cheap Chinese models to hedge the effect that OpenAI/Anthropic collapsing would have on their portfolios. This would make sense for (hedge funds holding) most companies: they don't really care about who supplies the AI, as long as they get it at roughly the same price as their competitors.

  • liuliu 59 minutes ago
    DeepSeek 4.1 Flash 0910 is perfect for M5 Ultra 256GiB. Running it fully resident in RAM, prefill at ~2500 tok/s and decode at ~40 tok/s. Probably tons of room to improve from there.
    • sdg03ksdv0d 49 minutes ago
      You are running this now? :o
  • smallmancontrov 1 hour ago
    They might be. They would delay public admission as long as possible, because public admission would make stocks go down.
  • xyzsparetimexyz 50 minutes ago
    There was a moment 3 months back where the sentiment was that cheaper models were the way to go. Since then the pendulum has swung back.
  • booi 1 hour ago
    Because GLM 5.3 Flash is even cheaper?
    • f311a 52 minutes ago
      Opencode Go gives only 6300 requests for glm and 23 000 for deepseek. And, if I wanted to, I would be able to do all my work on $10 plan with deepseek. It’s very cheap.
    • ActionHank 1 hour ago
      Nah fam, not true, also DS edges it out on coding / dev tasks.
      • jacquesm 1 hour ago
        That is opposite to my experience so far, can you describe your coding tasks? Mine are systems level code, utilities, operating system code, networking and real time control stuff.
        • UncleOxidant 44 minutes ago
          I also prefer GLM-5.3-flash to DS-4.1-flash, but it's close. Since Z.ai has been offering essentially free GLM-5.3-flash tokens on their coding plan between 8am-6pm pdt I've been using it a lot... though that ends on Oct 10 IIRC.
      • sampullman 1 hour ago
        On design tasks too, for me.
      • shellwizard 28 minutes ago
        [flagged]
    • wg0 1 hour ago
      Don't think so.
  • pianopatrick 1 hour ago
    I was just using a bunch of models in Cursor to review a project. I went looking for DeepSeek and it was not one of the options.

    Would be cool if they added it.

    • hirako2000 53 minutes ago
      Since they adhere to the same API spec, you can hook any model. It takes one line edit in /etc/hosts

      There are some quirks if your harness use unsupported features of course.

  • pizza234 40 minutes ago
    People have been raving since forever about Deepseek, but if one looks at the CoT, it's evident that it's way way stupider than frontier models (there's a reason why it's cheap). It's laughable to compare Deepseek 4.1 with Opus 5.5.

    I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).

    Local models are also really slow, unless one spends insane amounts of money.

    Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).

    • apitman 28 minutes ago
      The argument that most people are making isn't that dsv4.1f is better than frontier, but that it's good enough for most tasks, faster, and way cheaper.

      > if one looks at the CoT, it's evident that it's way way stupider than frontier models

      Frontier models don't show the full CoT

    • computerex 20 minutes ago
      The COT isn't an end all be all. Research has shown that the COT isn't necessarily what the model is actually thinking.
  • gsky 1 hour ago
    America bans Chinese models sooner or later just the China banned American big tech
  • tengbretson 1 hour ago
    I don't know about "freaking out", but I'd say I'm having a good time here with DS 4.1 flash.
  • hypfer 1 hour ago
    Is it known why unsloth seems to not have touched DeepSeek 4.1 Flash?
    • jacquesm 1 hour ago
      You can ask them directly, Daniel Han-Chen is pretty responsive.
  • cactusplant7374 22 minutes ago
    Because engineers are lusting for 1000 tokens per second. You can only achieve something like that with OpenAI.
  • kristianp 1 hour ago
    > shrank the KV cache by roughly 437X

    Can't you just say "shrank to 1/437th the size"? It's not that hard.

  • MisterMunchkin 1 hour ago
    I had it make 25 different things today and it cost $0.70

    It’s disgustingly good value. I find it capable of doing anything I want.

    Obviously can’t use it at work, but for home projects it’s awesome.

    • Octoth0rpe 41 minutes ago
      > Obviously can’t use it at work

      I do wonder how long it'll be before a us-hosted offering is available via bedrock, copilot, etc.

      • computerex 18 minutes ago
        There are already US hosted offerings on companies like fireworks.ai.
  • m3kw9 18 minutes ago
    i thought 6.1sol copied the caching architecture so this isn't such a big deal no more
  • pessimizer 41 minutes ago
    I'm no expert, but it think that it's the pricing on GPT-6 Luna. I'm also guessing that it's been underpriced just for this reason. I also don't think it's all that great, but it's definitely very cheap.

    If it's underpriced, it's a loss leader to sell the other models, so it actually can't be too good.

    I really put these things through their paces because I use them to review and work with new abstract game rules and models, so they're always flying blind. Luna misses the obvious (and more importantly, the clearly explained) consistently. My second prompt is listing all of the points in its first response, and saying "No, it doesn't work like that." The third prompt is picking out the two or three suggestions it made after correcting itself on all of the original points and saying "That's how it already works." The fourth prompt is "Now that we're done going over the rules, can we start?"

    I actually feel like 5.6 Luna seemed better.

  • anguralbanish2 1 hour ago
    I would love to get them more better, it's good not a bad thing.
  • sergiotapia 51 minutes ago
    In my experience it just takes so much longer to arrive at "done" state for me. It thinks for soooooo long. I guess if you're running 12 sessions at once you don't really notice.
  • doctorpangloss 59 minutes ago
    because it doesn't work very well?

    if you have a legitimate coding application, it isn't very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...

    • computerex 14 minutes ago
      What is your evidence? Deepseek v4.1 Flash is by far the most popular coding model on openrouter, having processed 38.7T tokens in just the last 7 days, over 3x the usage of the 2nd rank model.

      So I ask again, what are you basing your assertion on?

  • AIblemblio 1 hour ago
    No they can't.

    And as long as I pay as little for claude opus 5.5 i do right now, i'm using it.

    But yes i'm glad that we have alternatives.

  • verdverm 19 hours ago
    Why would we freak out? The systems we use have always gotten better, faster, cheaper with time
  • kydanet 41 minutes ago
    [flagged]
  • CurbStomper4 10 minutes ago
    [dead]
  • distantsounds 57 minutes ago
    because we've all figured out that AI is just a huge grift?
  • sroussey 1 hour ago
    Not comparing to gpt-6-luna which seems comparable and priced well.
  • wewewedxfgdf 55 minutes ago
    You might also choose to pay money for a service that provides real value instead of actively choosing to support the Chinese deliberate effort to undermine this country.
    • BarryMilo 47 minutes ago
      Your comment seems to imply there's a good guy in a this. I just see the inevitable end of an era, championed by predictably selfish actors.
    • f6v 12 minutes ago
      Oh, the writing is on the wall. Wait till you hear that European “sovereign AI” is just running GLM 5.3.
  • thefourthchime 1 hour ago
    For non-coding tasks it may be fine. But for coding, Opus 5.5 is just a completely another level than something like Deepseek 4.1 Flash.

    Opus 5.5: TIME 9.3m COST / $1.99 / SCORE 99/100 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5

    Deepseek 4.1 Flash: TIME 2.8m / COST $1.89 / SCORE 72/100 https://jonclegg.github.io/pacman-bakeoff/dev/#deepseek-v4.1...