OpenAI Jalapeño: Better than Nvidia Blackwell

(newsletter.semianalysis.com)

111 points | by bmulholland 5 hours ago

12 comments

  • jimmySixDOF 35 minutes ago
    I love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al
    • tmp10423288442 22 minutes ago
      SemiAnalysis’ founder was roommates with Anthropic people, not OpenAI, so he may be slightly (very slightly) more objective here.
    • xyzsparetimexyz 21 minutes ago
      s** posting? sex posting?
      • msh 9 minutes ago
        shit posting
    • verall 26 minutes ago
      semianalysis is pretty good
    • antonvs 19 minutes ago
      > I love how now you have to consider the possible s*** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods

      I mean, previously you could have said something much the same except substitute "frat boys".

  • fraboniface 16 minutes ago
    I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.
  • LarsDu88 15 minutes ago
    Well Sam Altman finally has built a moat against Chinese open weight AI. Well done. But what will this mean for Cerebras?

    I remember when Tesla was building its own inference chips, and after about 2 years and billions spent, the whole effort was scuttled b/c they simply could not keep up with the iteration and R&D cycles of dedicated chip companies. I suspect the same will be the case with OpenAI vs Cerebras + Nvidia/Groq

  • anthonypasq 49 minutes ago
    Continued hardware improvements really make it hard for me to believe token prices will not continue to plummet.
    • dgellow 28 minutes ago
      There is just so much downward pressure on token price, from every direction. We would need a completely new understanding of economics to explain why the price shouldn’t go down. Or market collusion/regulatory manipulation.
      • dumberquestions 3 minutes ago
        The demand for them is growing _per person_, not just across the wider economy, if tokens cost half as much but you want to use 3 times as much you're going to have to pay more.
      • simianwords 23 minutes ago
        The price has been going down for ages, its not clear what you are pointing at
    • jrflo 14 minutes ago
      This may just be a classic case of Jevons paradox: https://en.wikipedia.org/wiki/Jevons_paradox

      In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily.

      It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually just meant that we found more uses for steam engines.

      • kilroy123 12 minutes ago
        This is exactly what I see happening now.

        Codex keeps doing these usage resets. What do I do? Burn even more tokens than ever before. I know I'm not the only one.

      • anthonypasq 6 minutes ago
        the total cost spent on tokens may go up, but i just cant imagine per token costs going up
    • datakan 42 minutes ago
      Token prices coming down means nothing if the models keep wasting them
    • ilaksh 15 minutes ago
      Yeah but is it really even as good as Rubin? Seems just competitive.
    • mathisfun123 24 minutes ago
      this is a story about a proprietary accelerator being built/designed by a token provider. and you think they're going to return the efficiency gains to the customer instead of capture the value for themselves? interesting take.
      • anthonypasq 5 minutes ago
        OpenAI just dropped the price of Luna by 80% and Sol by 20-30%
        • mathisfun123 3 minutes ago
          and amazon shipping used to be free without prime, and uber used to be cheaper than taxis, and airbnb used to be cheaper than hotels.

          you really don't get it?

      • spacephysics 17 minutes ago
        We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss

        So as much as i agree “more profits to stakeholders screw the customer”, i think its more of an emergency to get to profitability before the music stops.

        • anthonypasq 5 minutes ago
          > We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss.

          what makes you think this?

      • simianwords 18 minutes ago
        Yes, I can bet on this happening. If anything, this is a net gain for consumers as it is a competitive market.
        • mathisfun123 2 minutes ago
          go ahead and bet: alibaba is a publicly traded company
    • gwerbin 33 minutes ago
      Hopefully this also means billionaires can stop trying to drop data centers into residential neighborhoods with zero noise control and polluting on-site generators, signing local politicians on with NDAs, calling for eminent domain to seize homes to build power lines to data centers, etc. etc. etc. Not to mention the water use controversy.

      Token prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.

      • tmp10423288442 20 minutes ago
        Nah, Jevon’s Paradox says that cheaper tokens will mean increased overall energy consumption.

        If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.

  • throwaw12 7 minutes ago
    Competition is good for all of us, we will get better and faster chips.

    Or at least Nvidia GPUs will become slightly cheaper for regular consumers again

  • thebeardisred 11 minutes ago
    All of these words spilled and no mention of the ISA.
  • ChoosesBarbecue 1 hour ago
    This is most impressive. The interesting question to me, is outside of the LLM accelerator space: will generalized chips have massive leaps in performance once LLM technology is used to create the next generation? In general, will we see rapid advances while we extract the value of these models in creating architectures? I'm so far removed from the space that this is a very naive interpretation of all this, but I'm curious.
  • epistasis 39 minutes ago
    It's so funny to see FP4.... I remember 20 years ago being asked what sort of HPC we needed in genomics, and the answer was basically, "lower precision, faster" for the stuff I was working on. But FP4 is, well, almost comical.

    One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)

    • nxtfari 31 minutes ago
      Agree, I remember when even half precision made its way into C# sometime around 2020 (I didn’t know much about ML then) and I thought, well I guess that’s a worthwhile tradeoff but I can’t imagine going lower. Lo and behold (1-bit Bonsai) how much lower you could go.
  • empath75 23 minutes ago
    When people talk about the commodification of inferencing, they imagine a future where everyone has access to frontier models and can run them at the same cost, and what will actually happen is closer to the commodification of _oil_, where only a few companies have the scale to produce it at a competitive price, and advances like this are _why_.

    Once models are more or less interchangeable, the price of LLMs will drop to essentially the price of energy required to run them, and the big labs will be able to run them cheaper than anyone else.

    • simianwords 18 minutes ago
      I don't believe models will be commodified because each model is unique with strengths and weaknesses. Its not like Steel which is more or less the same no matter where you purchase it from.

      If what you said were true, you would hardly see people complaining about the quality of Opus 5 or good writing from Sol. But people do.

      • airspresso 5 minutes ago
        This depends heavily on what the use-case is. Yes, if it's a coder making software and having to read LLM output then writing style matters. If the LLM is used in an automated data processing pipeline with a capped level of complexity, entirely different aspects matter and LLMs become more interchangeable.
  • simianwords 12 minutes ago
    How can OpenAI mass produce this chip at scale more economically than Nvidia which has experience in the supply chain and scale efficiencies to do it efficiently?
    • dpe82 0 minutes ago
      NVidia has enormous operating margins, so a competitive solution doesn't have to match or beat NVidia's scale efficiencies; it just has to beat delivered cost.

      One objective of the project might be simply to provide credible negotiating leverage when dealing with existing suppliers like NVidia. You don't have to deploy at scale for that to work, but you do have to look like you could if pushed hard enough.

    • chris_money202 6 minutes ago
      In the short and medium term, it probably won't be more economical to produce for OpenAI. Where OpenAI is benefitting from their own chip is being able to tailor it to their models and workloads. When you buy off the shelf Nvidia, its not perfectly tailored and OpenAI has to spend marginally more to run off that chip. At the scale OpenAI is operating at and plans to operate at, that margin becomes pretty big $$
    • airspresso 8 minutes ago
      By leveraging the experience Broadcom has in this area. Still remains to be seen how that goes when they want to scale production.
  • 0xbadcafebee 8 minutes ago
    [delayed]
  • varispeed 45 minutes ago
    Why they don't research how to make their own RAM and they have to buy it from the common market?

    They should GTFO with this crap.

    Create barriers to computing for ordinary people while milking businesses for tokens.

    • petcat 40 minutes ago
      Building a custom-designed ASIC is much easier than producing state of the art memory chips.

      There's a reason why Micron and Nvidia are the crown jewels of American technology right now and for the foreseeable future.

      • JV00 36 minutes ago
        Nvidia does not make RAM
      • brcmthrowaway 37 minutes ago
        NVIDIA produces memory?
        • fc417fc802 30 minutes ago
          Fabless AFAIK. And that's the actual problem - drawing up CAD diagrams doesn't help if the factories are fully booked out.
        • Cyph0n 18 minutes ago
          A state of the art GPU is much harder to design & produce at scale and than an internal ASIC.
    • datakan 39 minutes ago
      People keep saying stuff like this without understanding what it takes to make RAM. It's one of, if not the most, heavily patented things in the world. The second you dip your toes into those waters the lawsuits begin.

      If somehow you get around the patent issues, you're now faced with huge research and development costs, fabs to build, processes to sort out and all of that has very high failure rates.

      Last time I checked Micron was the largest patent holder in the world and even for them this is a hard area where they are number 3 in the market.

      • chris_money202 2 minutes ago
        RAM chips are not hard to produce compared to many other types of semiconductors; Intel started in the memory game and left because the margins weren't great and they were going to fold. The failure rates on these chips are actually very tolerable; you can have a very bad yield and still have a viable chip due to things like ECC.