시장보고서
상품코드
2100786

2026년 동향 : 새로운 AI 추론 수요에 대응하는 메모리

2026 Trends: Memory for New AI Inference Demand

발행일: | 리서치사: 구분자 TrendForce | 페이지 정보: 영문 11 Pages | 배송안내 : 1-2일 (영업일 기준)

    
    
    



가격
PDF (Corporate License) help
PDF 보고서를 동일 사업장에서 5명까지 이용할 수 있는 라이선스입니다. 인쇄 가능하며 인쇄물의 이용 범위는 PDF 이용 범위와 동일합니다.
US $ 2,500 금액 안내 화살표 ₩ 3,594,000
※ 부가세 별도
한글목차
영문목차
※ 본 상품은 영문 자료로 한글과 영문 목차에 불일치하는 내용이 있을 경우 영문을 우선합니다. 정확한 검토를 위해 영문 목차를 참고해주시기 바랍니다.

2026년 1월, NVIDIA는 BlueField-4 DPU가 관리하는 'CMX 컨텍스트 메모리 스토리지 플랫폼'을 발표했습니다. 이는 로컬 SSD와 공유 스토리지 간의 메모리 계층을 확장하여, AI 추론 시대에 발생하는 방대한 KV 캐시 스토리지 수요에 대응하기 위한 것입니다. 또한, NVIDIA와 Arm은 에이전트형 AI의 CPU 요구 사항을 충족하기 위해 잇달아 CPU 랙을 출시하고 있으며, 이를 통해 CPU RAM의 새로운 시장이 창출되고 있습니다.

본 보고서에서는 (1) AI 추론에서의 메모리 수요, (2) KV 캐시 오프로드에 의해 주도되는 SSD POD 수요, 그리고 (3) 에이전트형 AI에 의해 주도되는 CPU RAM 수요에 대해 상세히 분석합니다. 그 목적은 AI 추론 시대에 메모리 용량 수요가 확대되고 있는 이유를 설명하고, 현재의 솔루션을 검증함과 동시에 향후 나타날 메모리 수요의 구조를 개괄하는 데 있습니다.

주요 하이라이트

  • NVIDIA는 AI 추론에서 대규모 KV 캐시를 위해 로컬 SSD와 공유 스토리지 사이를 연결하는 BlueField-4 DPU로 관리되는 CMX 플랫폼을 도입했습니다.
  • NVIDIA와 Arm은 에이전트형 AI의 CPU 수요를 충족하기 위해 CPU 랙을 출시했으며, 이로 인해 CPU 메모리 수요가 단계적으로 증가하고 있습니다.
  • 본 보고서에서는 AI 추론에서의 메모리 수요, KV 캐시 오프로드에 따른 SSD POD 수요, 그리고 에이전트형 AI로 인해 유발되는 CPU 메모리 구조의 변화에 초점을 맞추고 있습니다.

목차

  • 1. AI 추론에서의 메모리 수요
  • 2. SSD POD 수요는 KV 캐시 오프로드에 의해 촉진
  • 3. 에이전트형 AI에 따른 CPU 수요 증가
  • 4. TRI의 견해
KSM 26.08.03

In January 2026, NVIDIA introduced the CMX Context Memory Storage Platform, managed by the BlueField‑4 DPU, to extend the memory hierarchy between local SSD and shared storage and address the massive KV cache storage demands of the AI inference era. In addition, NVIDIA and Arm have successively launched CPU racks to meet the CPU requirements of agentic AI, creating an incremental market for CPU RAM.

This report provides an in‑depth analysis of: (1) memory demand in AI inference; (2) SSD POD demand driven by KV cache offloading; and (3) CPU RAM demand driven by agentic AI. The goal is to explain why memory capacity needs are expanding in the AI inference era, review current solutions, and outline the future structure of emerging memory demand.

Key Highlights

  • NVIDIA introduced the CMX platform managed by BlueField‑4 DPU to extend between local SSD and shared storage for large KV cache in AI inference.
  • NVIDIA and Arm launched CPU racks to meet agentic AI CPU needs, creating incremental CPU memory demand.
  • The report focuses on AI inference memory needs, SSD POD demand from KV cache offload, and CPU memory structure changes driven by agentic AI.

Table of Contents

  • 1. Memory Demand in AI Inference
    • Figure 1: AI Models Average Output Tokens per Question (2023-2026)
    • Figure 2: Example of KV Cache Applications
    • Figure 3: Changes to CPU:GPU Ratio among Agentic AI Applications
  • 2. SSD POD Demand Driven by KV Cache Offloading
    • Figure 4: Sequence of KV Cache Offloading for NVIDIA’s Dynamo (G1-G4)
  • 3. CPU Demand Driven by Agentic AI
    • Figure 5: NVIDIA’s Vera CPU Architecture
    • Table 1: CPU Specifications of Various Suppliers (2023-2026)
    • Table 2: Analysis on Hypothetical Shipment Scenario of NVIDIA’s CPUs in 2026
    • Figure 6: Analysis Results on Demand Scenario of NVIDIA’s CPUs in 2026
    • Table 3: Summary of Memory Demand Drivers Introduced by AI Inference
  • 4. TRI’s View
샘플 요청 목록
0 건의 상품을 선택 중
목록 보기
전체삭제
문의
원하시는 정보를
찾아 드릴까요?
문의주시면 필요한 정보를
신속하게 찾아드릴게요.
02-2025-2992
email
문의하기