Ask HN: Can token caching be a business idea?

1 points | by sourav_biswas 3 hours ago

1 comments

  • minimaxir 3 hours ago
    No, because the cache is local to the GPUs and there would be no cost/compute reason to have long-lived caches outside of the 1 hour cache already done by LLM APIs.
    • sourav_biswas 3 hours ago
      What if the real business isn't caching generic prompts/but caching outputs for a specific vertical where one can know two requests are truly equivalent n not just embedding similar?