No, because the cache is local to the GPUs and there would be no cost/compute reason to have long-lived caches outside of the 1 hour cache already done by LLM APIs.
What if the real business isn't caching generic prompts/but caching outputs for a specific vertical where one can know two requests are truly equivalent n not just embedding similar?
1 comments