{"id":43907,"date":"2026-07-23T23:30:00","date_gmt":"2026-07-23T21:30:00","guid":{"rendered":"http:\/\/stocks-future.com\/?guid=fa502839b97370c4dc41d23a274a3b46"},"modified":"2026-07-23T23:30:00","modified_gmt":"2026-07-23T21:30:00","slug":"tensormesh-and-amd-collaborate-to-empower-fewer-gpus-to-serve-more-models","status":"publish","type":"post","link":"https:\/\/stocks-future.com\/?p=43907","title":{"rendered":"Tensormesh and AMD Collaborate to Empower Fewer GPUs to Serve More Models"},"content":{"rendered":"<p class=\"bwalignc\">LMCache enables AMD GPUs to overflow KV cache onto a multi-tier memory hierarchy spanning DRAM, SSD and remote storage<\/p><p>SAN FRANCISCO--(BUSINESS WIRE)--<a  href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fwww.tensormesh.ai%2F&amp;esheet=54575412&amp;newsitemid=20260723669125&amp;lan=en-US&amp;anchor=Tensormesh&amp;index=1&amp;md5=52413822f98dd4ec19d502dc55a4762b\" rel=\"nofollow\" shape=\"rect\">Tensormesh<\/a>, the company pioneering caching-accelerated inference optimization for enterprise AI, today announced a collaboration with AMD through which Tensormesh KV cache solution and AMD virtual memory offering will work together to allow more models to be served on fewer GPUs while retaining high KV cache hit rates and throughput even with oversubscribed high-bandwidth memory (HBM). Tensormesh is working with AMD, leveraging its GPU technology and AMD Live Context Virtualization components, and is tested using Dell servers with 8x AMD\/ATI accelerators (MI355) GPUs and Dell storage. Tensormesh integrated LMCache coordinates KV cache management.<\/p><br\/><a href=\"https:\/\/mms.businesswire.com\/media\/20260723669125\/en\/2858209\/5\/Tensormesh_Logo.jpg\"><img src=\"https:\/\/mms.businesswire.com\/media\/20260723669125\/en\/2858209\/22\/Tensormesh_Logo.jpg\" \/><\/a><br\/><a href=\"https:\/\/mms.businesswire.com\/media\/20260723669125\/en\/2858209\/5\/Tensormesh_Logo.jpg\"><img src=\"https:\/\/mms.businesswire.com\/media\/20260723669125\/en\/2858209\/21\/Tensormesh_Logo.jpg\" \/><\/a><p>For users, running more models on the same set of GPUs means lower costs. LMCache users can now reuse the infrastructure they have built to virtually expand their GPUs' capacity. In addition, customers can reuse KV cache chunks stored for short-term memory virtualization later for prefix or non-prefix KV cache matching, maximizing system efficiency.<\/p><p><b>The Results<\/b><\/p><p>This new approach is far more efficient than building GPUs with more memory, which has led to the industry\u2019s current memory supply crisis and increased the number of GPUs. Enterprises running AI over large document sets can now have a better experience and a lower bill. Testing showed:<\/p><ul class=\"bwlistdisc\"><li><b>Near-instant responses at scale: <\/b>On a 300 GB document set (Kimi-K2.6), time-to-first-token dropped from 3.4 seconds to under half a second, a nearly 7x improvement, thanks to reusing cached KV data from DRAM and NFS instead of recomputing it from scratch.<\/li><li><b>No slowdown as usage grows: <\/b>Output throughput held steady at ~48 tokens per second regardless of workload size, while unoptimized inference throughput fell by nearly 40% under the same load.<\/li><li><b>Twice the model density:<\/b> The same hardware doubled the model density.<\/li><\/ul><p>\u201cThis powerful new collaboration builds on <a  href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fwww.businesswire.com%2Fnews%2Fhome%2F20260527958597%2Fen%2FTensormesh-Raises-%252420M-from-Investors-Including-AMD-Ventures-CoreWeave-NVentures-Launches-Tensormesh-Inference-to-Fix-AIs-Most-Expensive-Problem&amp;esheet=54575412&amp;newsitemid=20260723669125&amp;lan=en-US&amp;anchor=AMD%27s+recent+strategic+investment+in+Tensormesh&amp;index=2&amp;md5=48bc6a7058b07846e677bbed4c9a5c54\" rel=\"nofollow\" shape=\"rect\">AMD\u2019s recent strategic investment in Tensormesh<\/a> and expands the capabilities of AMD GPUs,\u201d explained Junchen Jiang, Tensormesh CEO and LMCache co-creator. \u201cTogether, we\u2019re greatly enhancing the memory that inference engines can access for model weights and KV cache, using all of the memory resources on each node.\u201d<\/p><p>\u201cWe recognize Tensormesh and LMCache as KV cache management leaders,\u201d said Anush Elangovan, AMD\u2019s vice president of AI software. \u201cAnd we\u2019re thrilled to announce AMD\u2019s breakthrough in virtual GPU memory management, amplified by LMCache and Tensormesh.\u201d<\/p><p>The companies are sharing performance results from their collaboration at this week\u2019s <a  href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fwww.amd.com%2Fen%2Fcorporate%2Fevents%2Fadvancing-ai.html&amp;esheet=54575412&amp;newsitemid=20260723669125&amp;lan=en-US&amp;anchor=AMD+Advancing+AI+2026&amp;index=3&amp;md5=d56b5961729c8b9be6df02ddcaac2fa0\" rel=\"nofollow\" shape=\"rect\">AMD Advancing AI 2026<\/a>, a month before the public release of the LMCache open-source code.<\/p><p><b>Additional Resources:<\/b><\/p><ul class=\"bwlistdisc\"><li><a  href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fwww.tensormesh.ai%2Fblog-posts%2Ffixing-ais-most-expensive-problem----junchen-jiang-tensormesh-ceo&amp;esheet=54575412&amp;newsitemid=20260723669125&amp;lan=en-US&amp;anchor=Fixing+AI%27s+Most+Expensive+Problem+with+Junchen+Jiang%2C+Tensormesh+CEO&amp;index=4&amp;md5=6ffd27a7ed70644a55267f7014454b28\" rel=\"nofollow\" shape=\"rect\">Fixing AI\u2019s Most Expensive Problem with Junchen Jiang, Tensormesh CEO<\/a><\/li><li><a  href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Flmcache.ai%2F&amp;esheet=54575412&amp;newsitemid=20260723669125&amp;lan=en-US&amp;anchor=LMCache%3A+Build+The+Foundation+of+AI+Memory+Tensor+with+KV+Cache+Infrastructure&amp;index=5&amp;md5=bc9d61074b0e775a339e7b193a8dead2\" rel=\"nofollow\" shape=\"rect\">LMCache: Build The Foundation of AI Memory Tensor with KV Cache Infrastructure<\/a><\/li><li><a  href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fwww.tensormesh.ai%2Fblog-posts%2Fpersistent-kv-cache-for-ai-inference&amp;esheet=54575412&amp;newsitemid=20260723669125&amp;lan=en-US&amp;anchor=Persistent+KV+Cache%3A+Own+Your+Context+Caching+Lifecycle&amp;index=6&amp;md5=0940558b75f47b6f4da802605a0eb8be\" rel=\"nofollow\" shape=\"rect\">Persistent KV Cache: Own Your Context Caching Lifecycle<\/a><\/li><\/ul><p><b>About Tensormesh<\/b><\/p><p>Tensormesh is the leader in caching-accelerated inference optimization for enterprise AI. Founded by faculty, PhD researchers and alumni from the University of Chicago, UC Berkeley, and Carnegie Mellon, and led by Junchen Jiang, University of Chicago faculty member and co-creator of LMCache, Tensormesh builds on years of academic research in distributed systems and AI infrastructure. The company has raised $24.5 million in total funding and is backed by Valley Capital Partners, NVentures, AMD Ventures, CoreWeave, and Laude Ventures.<\/p><br\/> <b>Contacts<\/b> <br\/><p><b>Media Contact<\/b><br\/><a  href=\"mailto:press@tensormesh.ai\" rel=\"nofollow\" shape=\"rect\">press@tensormesh.ai<\/a><\/p>","protected":false},"excerpt":{"rendered":"<p>LMCache enables AMD GPUs to overflow KV cache onto a multi-tier memory hierarchy spanning DRAM, SSD and remote storageSAN FRANCISCO&#8211;(BUSINESS WIRE)&#8211;Tensormesh, the company pioneering caching-accelerated inference optimization for enterprise AI, today&#8230;<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-43907","post","type-post","status-publish","format-standard","hentry","category-infos-businesswire"],"_links":{"self":[{"href":"https:\/\/stocks-future.com\/index.php?rest_route=\/wp\/v2\/posts\/43907","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/stocks-future.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/stocks-future.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/stocks-future.com\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/stocks-future.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=43907"}],"version-history":[{"count":1,"href":"https:\/\/stocks-future.com\/index.php?rest_route=\/wp\/v2\/posts\/43907\/revisions"}],"predecessor-version":[{"id":43908,"href":"https:\/\/stocks-future.com\/index.php?rest_route=\/wp\/v2\/posts\/43907\/revisions\/43908"}],"wp:attachment":[{"href":"https:\/\/stocks-future.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=43907"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/stocks-future.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=43907"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/stocks-future.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=43907"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}