NVIDIA AI Infrastructure’s cover photo
NVIDIA AI Infrastructure

NVIDIA AI Infrastructure

Computer Hardware Manufacturing

Santa Clara, California 302,284 followers

About us

The NVIDIA accelerated computing platform is optimized for energy efficiency while accelerating AI performance, helping enterprises deploy secure, future-ready AI data centers.

Website
https://www.nvidia.com/en-us/data-center/
Industry
Computer Hardware Manufacturing
Company size
10,001+ employees
Headquarters
Santa Clara, California

Updates

  • 💡 NVIDIA Vera is the CPU for agents and benchmark results from DeepInfra show that it's more than 2x as fast compared with other CPUs. Read the results ⤵️

    View organization page for DeepInfra

    3,835 followers

    🚀 We run AI agents in production so when NVIDIA built Vera CPU specifically for agents, we measured it ourselves rather than relying on the spec sheet. NVIDIA Vera CPU(88 Olympus cores, the CPU side of the upcoming Vera Rubin platform) is designed around a real shift: agentic AI turns inference into a loop, and the CPU executes all the work around each model call orchestration, parsing, sandboxed code execution, data movement. When the CPU is slow, the agent waits. We got early access to Vera and tested it with our own harness: a real agentic workload from our production environment, real traffic replayed at fixed pacing (zero live tokens, so every difference is the CPU), and an identical fair configuration across four architectures NVIDIA Vera, AMD Zen5, Intel Granite Rapids, Intel Sapphire Rapids. What we found: → NVIDIA's public claim is “80% faster agentic CPU performance” (a 1.8× bar). We locked our comparison basis before measuring Vera. Measured result: 1.8x on orchestration over best x86 CPU and the fastest of all four architectures tested. → Vera swept every workload category, orchestration, code-execution spawn, memory bandwidth, and base throughput where no single incumbent had dominated. → At a strict QoS bar, Vera sustained 256 concurrent agents per partition vs 160 for incumbents: 1.33–1.6× more agents per core. → With the agent fleet at full load, we served a real LLM on the 56 leftover cores at 118.5 tok/s, more than a full 72-core Grace socket doing nothing else. The mechanism is real per-core execution: 3.62 IPC on branch-heavy work versus 2.44 for the best measured x86 (Intel; AMD Zen5 had no IPC capture) Full technical write-up, including methodology: → https://lnkd.in/g9ryzh_8

  • As AI workloads continue to grow, peak compute alone doesn't cut it anymore. High-bandwidth, low-latency, purpose-built scale-up networking is critical to AI factory performance. That's the job of #NVIDIANVLink — the scale-up network purpose-built for AI factories. Designed so every chip in the system works as one, NVLink delivers the bandwidth, low latency, and intelligent resiliency that production AI infrastructure depends on. From the chip level to the software stack, it's all co-designed to move at the speed of AI innovation. Read the full breakdown: https://nvda.ws/4wmHIw4

    • No alternative text description for this image
  • We built NVIDIA Confidential Computing so regulated industries can securely run AI on their most sensitive data. 🛡️ Using hardware-enforced Trusted Execution Environments, enterprises get cryptographic proof that their sensitive data and AI models are protected throughout every stage of inference — from the model owner down to the infrastructure operator. Supported across NVIDIA Hopper, NVIDIA Blackwell, and NVIDIA Vera Rubin platforms. Learn more about confidential computing 👉 https://nvda.ws/4wRykAb

  • NVIDIA AI Infrastructure reposted this

    🚀 Bristol Myers Squibb is fundamentally changing how we approach drug discovery. They’ve just announced the deployment of their second AI factory - the most powerful in the life sciences industry, built on NVIDIA Vera Rubin architecture. By integrating eight NVIDIA DGX Vera Rubin NVL72 systems, BMS is achieving up to 10x the performance per megawatt compared to previous systems. A game-changer for healthcare: 🧬 Democratized Compute: Opening unified supercomputing access to BMS scientists globally 🤖 Agentic AI: Utilizing the NVIDIA BioNeMo Agent Toolkit for complex biological predictions and model training ⚡ Streamlined Discovery: Eliminating traditional research bottlenecks to drastically accelerate the journey from lab to market We are moving past isolated AI projects and entering an era where AI factories are the backbone of biomedical research. 📃https://nvda.ws/4vAoMsh

    • No alternative text description for this image
  • More quantum chemistry. More GPU acceleration. More possibilities for discovery. ⚡ 📣 NVIDIA cuEST 0.2.0 is available now, expanding GPU-accelerated electronic structure calculations with: 🔬 Density-fitted LRC exchange (ωK) support for symmetric matrices and gradients ⚛️ Density-fitted nonsymmetric exchange (K + ωK) matrix 🚀 AO-to-MO integral tensor transformation in the density fitting representation 🧪 Extended PCM derivatives support and more Learn more now ➡️ https://nvda.ws/4fqYP8H

    • No alternative text description for this image
  • 📣 We're collaborating with SoftBank R&D on our next phase of work on AI-native networks and physical AI for Japan. Our CEO Jensen Huang met with SoftBank Corp. president and CEO Junichi Miyakawa to discuss what's next. SoftBank is using NVIDIA GB200-class infrastructure, RTX PRO-based AI-RAN with NVIDIA AI Aerial, and Nemotron-based large telecom models to turn its communications network into an intelligence delivery network. 🔗 https://nvda.ws/4vzGoVi

    • No alternative text description for this image
  • The #NVIDIAVeraRubin platform consists of five rack-scale systems for AI agents with a supply chain spanning over 350 factory sites across 30 countries with millions square feet of factory floor space. Vera Rubin is in full production, co-designed with NVIDIA DSX and DSX MaxLPS to deliver the lowest token cost and maximum tokens per watt. Congratulations to our partners who have their engineering racks up and running: CoreWeave, Dell Technologies, Microsoft, and Oracle Cloud.

  • Our internal AI factory serves 4 trillion tokens a month to NVIDIA employees, and demand is still growing 40% month over month. In Episode 2 of AI Factory Insider, we get into what it actually takes to run that at scale: the architecture, the real use cases, and what we learned along the way. Catch the full episode ➡️ https://nvda.ws/3T4nPLn

  • Perplexity launched SPACE, a secure sandbox platform built for agentic AI. Early tests on NVIDIA Vera CPU showed up to 1.9x faster sandbox starts. Faster starts = less latency, more parallelism, and agents that scale. Learn more now ⬇️

    View organization page for Perplexity

    1,718,606 followers

    Introducing SPACE, the sandbox platform behind Perplexity Computer. It creates isolated environments for code, files, and long-running agent sessions. SPACE has handled 100% of Computer production traffic since June. https://lnkd.in/gxNTgx-Q Agent infrastructure must be functional, efficient, and secure and traditional sandboxes were built for short-lived code execution. Agents need to run code, edit files, and run for hours or days. Runtimes must preserve work without leaving credentials inside environments. SPACE separates the session from the sandbox running it. Each task gets a disposable Firecracker microVM that is destroyed when the work ends. Rolling snapshots preserve live memory and files, so the session can pause, resume, or branch across sandboxes. On identical production traffic, SPACE reduced median sandbox creation latency from 185 ms to 60 ms. P90 fell from 447 ms to 89 ms. Last week it handled millions of sandbox creations and tens of millions of reconnects for Computer. Read more: https://lnkd.in/gSFrn7bW

    • No alternative text description for this image
    • No alternative text description for this image

Affiliated pages

Similar pages