logo
Αρχική Σελίδα Ειδήσεις

εταιρικά νέα για Inferra is Lightbits inferencing memory wall buster

Πιστοποίηση
Κίνα Beijing Qianxing Jietong Technology Co., Ltd. Πιστοποιήσεις
Κίνα Beijing Qianxing Jietong Technology Co., Ltd. Πιστοποιήσεις
Αναθεωρήσεις πελατών
Το προσωπικό πωλήσεων της Co. τεχνολογίας του Πεκίνου Qianxing Jietong, ΕΠΕ είναι πολύ επαγγελματικό και υπομονετικό. Μπορούν να παρέχουν τις αναφορές γρήγορα. Η ποιότητα και η συσκευασία των προϊόντων είναι επίσης πολύ υψηλές. Η συνεργασία μας είναι πολύ ομαλή.

—— 《Festfing DV》 LLC

Όταν έψαχνα τη Intel ΚΜΕ και Toshiba SSD επειγόντως, αμμώδης από το Πεκίνο Qianxing Jietong η Co. τεχνολογίας, ΕΠΕ μου έδωσε πολλή βοήθεια και με πήρε τα προϊόντα που χρειάστηκα γρήγορα. Την εκτιμώ πραγματικά.

—— Γεν γατακιών

Αμμώδης του Πεκίνου Qianxing Jietong η Co. τεχνολογίας, ΕΠΕ είναι πολύ προσεκτικός πωλητής, ο οποίος μπορεί να υπενθυμίσει σε με τα λάθη διαμόρφωσης εγκαίρως πότε αγοράζω έναν κεντρικό υπολογιστή. Οι μηχανικοί είναι επίσης πολύ επαγγελματικοί και μπορούν γρήγορα να ολοκληρώσουν την εξεταστική διαδικασία.

—— Strelkin Mikhail Vladimirovich

Είμαστε πολύ ευχαριστημένοι με την εμπειρία μας συνεργασίας με την Beijing Qianxing Jietong. Η ποιότητα των προϊόντων είναι εξαιρετική και η παράδοση γίνεται πάντα στην ώρα της. Η ομάδα πωλήσεων είναι επαγγελματική, υπομονετική και πολύ εξυπηρετική με όλα μας τα ερωτήματα. Εκτιμούμε πραγματικά την υποστήριξή τους και προσβλέπουμε σε μια μακροχρόνια συνεργασία. Συνιστάται ανεπιφύλακτα!

—— Ahmad Navid

Ποιότητα: Μεγάλη εμπειρία με τον προμηθευτή μου.Το MikroTik RB3011 είχε ήδη χρησιμοποιηθεί, αλλά ήταν σε πολύ καλή κατάσταση και όλα λειτουργούν τέλεια.Η επικοινωνία ήταν γρήγορη και ομαλή.Και όλες μου οι ανησυχίες λύθηκαν γρήγορα.- Πολύ αξιόπιστος προμηθευτής.

—— Γκεράν Κολέσιο

Είμαι Online Chat Now
επιχείρηση Ειδήσεις
Inferra is Lightbits inferencing memory wall buster


I’ve compressed the full text by 8% word count, retained all core facts, benchmarks, quotes and technical logic, and removed all links:

Lightbits’ Inferra software accelerates KV cache operations, breaks AI inference’s memory wall and drastically improves GPU utilization. The high-performance block storage firm built Inferra to virtualize GPU high-bandwidth memory (HBM). A KV cache stores pre-computed attention keys and values during LLM inference, avoiding redundant recomputation of sequence data and boosting generative AI performance.


Standard KV caching confines cached data to GPU HBM. Once HBM saturates, existing KV pairs are evicted for new data, requiring time-consuming recomputation upon reuse. Modern tiered KV caching extends storage to server DRAM, local NVMe SSDs and remote networked NVMe drives. While higher tiers incur longer access latency, this delay is often shorter than recomputing extensive token vectors.


Inferra addresses critical memory bottlenecks stemming from expanding long-context KV cache footprints. It enables cloud and enterprise users to run faster, more scalable AI inference without costly GPU and HBM hardware upgrades to expand cache capacity.


τα τελευταία νέα της εταιρείας για Inferra is Lightbits inferencing memory wall buster  0


Lightbits Labs Co-Founder and Chairman Avigdor Willenz noted the $117 billion 2026 AI inference market forces legacy training-focused vendors to re-architect. He stressed legacy storage and training systems cannot be retrofitted for KV cache bottlenecks, as they were designed for outdated workload paradigms. Purpose-built for GPU inference efficiency, Inferra has validated value in customer beta programs, extending the vendor’s NVMe over TCP innovation to solve modern inference pain points.


Powered by KV-cache-optimized prefetching algorithms, Inferra proactively loads upcoming KV cache data before GPU demand. It delivers three core advantages: up to 16× more concurrent inference sessions on existing GPU hardware with strict SLA compliance (no hardware upgrades required); up to 100× latency improvement and support for 10-million-token context windows, cutting TTFT and TPOT via predictive attention state prefetching instead of recomputation; and secure tenant isolation with intelligent unified KV cache management for shared inference clusters.


Inferra redefines inference economics by virtualizing GPU memory across multi-tier memory and storage, turning fragmented KV caching into a persistent, intelligent data layer. Deployed on servers’ x86 subsystems, it monitors KV cache block activity and leverages a Sub-Linear Sparse Attention Prefetch (SLSAP) engine. Combining locality-sensitive hashing and historical attention reuse patterns, it ranks and prioritizes frequently reused KV blocks, including recent tokens, semantically similar content and structural patterns in multi-turn chat and RAG workflows. The system retrieves targeted blocks from server DRAM or external SSDs and preloads them to GPU HBM via RDMA.


OVH Chief Product & Technology Officer Yaniv Fdida verified Inferra delivers substantial GPU utilization gains, enabling scalable, cost-efficient AI agent and RAG workload infrastructure.


τα τελευταία νέα της εταιρείας για Inferra is Lightbits inferencing memory wall buster  1


Public benchmark results confirm drastic performance lifts: Qwen 2.5-7B (410K context) speeds up 102× (72.6s → 711ms); DeepSeek-R1-70B (141K context) improves 152× (70.8s → 465ms); Llama-4-Scout (10M context) achieves a 380× leap (1.5 hours → 13s).


Hardware-agnostic and compatible with all mainstream open-source inference frameworks, Inferra supports Dynamo, vLLM and LMCache plugins. Uniquely, it delivers sub-linear TTFT scaling relative to context length, while competitors including DDN, VAST Data and WEKA only offer linear scaling or reactive streaming. Lightbits has demonstrated stable 10M-token context inference on commodity L40S GPUs, a benchmark unmatched by rival vendors.


Lightbits Labs AI Product and Business SVP Ramesh Chettuvetty stated Inferra delivers immediate TCO savings, eliminating long-context inference stalls and empowering businesses to run larger models and longer dialogue workloads at lower infrastructure costs. Users can assess Inferra’s inference cost optimization benefits via the vendor’s public efficiency evaluation platform. In a joint March demo with ScaleFlux, Inferra achieved 100× to 280× KV cache acceleration on computational storage SSD workloads.


Beijing Qianxing Jietong Technology Co., Ltd.
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!


Χρόνος μπαρ : 2026-09-10 14:12:24 >> κατάλογος ειδήσεων
Στοιχεία επικοινωνίας
Beijing Qianxing Jietong Technology Co., Ltd.

Υπεύθυνος Επικοινωνίας: Ms. Sandy Yang

Τηλ.:: 13426366826

Στείλετε το ερώτημά σας απευθείας σε εμάς (0 / 3000)