|
시장보고서
상품코드
2129080
자동차 및 로봇 VLA 대규모 모델 용도 조사 보고서(2026년)Research Report on Application of VLA Large Model in Automobiles and Robots, 2026 |
||||||
자동차 및 로봇 분야에서의 VLA 조사 : 하이브리드 아키텍처가 주류를 이루며, VLA는 일반적인 세계 모델과 통합되고, 강화 학습이 핵심 엔진으로 기능하고 있습니다.
VLA 모델이란 시각, 언어, 동작의 각 모달리티를 통합한 모델입니다. 통합된 멀티모달 학습 프레임워크를 채택하고, 지각, 추론, 제어를 통합함으로써 시각 입력(이미지·동영상)이나 언어적 지시로부터 실행 가능한 물리적 세계의 액션(예 : 로봇의 관절 운동, 차량의 조향·가속·제동 제어 등)을 직접 생성합니다.
VLA는 VLM(이해와 기술) + E2E(엔드투엔드 의사결정 및 제어) + CoT(인간과 같은 추론의 연쇄)에 해당합니다. 기능 비교의 관점에서 볼 때, VLA는 정밀한 3D 지각, 상식적인 이해, 논리적 사고 및 해석 가능성을 동시에 실현합니다.
자율주행 분야에서 VLA의 진화는 4단계로 나눌 수 있습니다.
언어를 인터프리터로 활용하는 단계(Pre-VLA) : 언어 모델은 제어에 관여하지 않고, 장면의 기술만을 생성합니다.
모듈형 VLA : 언어는 의사결정을 위한 계획 구성 요소로 기능하지만, 다단계 워크플로우로 인해 지연이 발생합니다.
통합형 엔드투엔드 VLA : 센서 입력이 단일 전방 전파를 통해 액션에 직접 매핑됩니다.
추론 기능이 강화된 VLA : LLM이 제어의 폐쇄 루프에 진입하여 장기적인 추론, 기억 및 대화 기능을 실현합니다. Li Auto의 MindVLA가 대표적인 예입니다.
구현 일정 :
분할형 엔드투엔드 방식은 2024년부터 2025년에 걸쳐 양산되었습니다.
단일 모델을 기반으로 한 엔드투엔드 및 VLA는 2025년부터 2026년에 걸쳐 광범위하게 도입되었습니다.
2026년은 자율주행용 대규모 모델에 있어 ‘다각적인 경쟁 심화와 패러다임 통합 가속화’라는 중요한 전환점이 될 것입니다.
기술적 현황 및 주요 지표.
광범위한 모델 파라미터 : NVIDIA Alpamayo 1.5에는 0.5B/10B의 파라미터가 포함되어 있습니다. DeepRoute.ai는 40B 파라미터의 기반 모델을 사용하고 있습니다. Afari Technology의 StepVL 기반 모델은 32B 파라미터를 특징으로 합니다(7B 및 3.6B로 증류); Li Auto의 MindVLA 32B 매개변수 기반 모델은 3.6B 매개변수의 MoE 변형으로 증류되었으며, 차량 모델에는 4B가 할당되어 있습니다.
차량용 실시간 성능 : Li Auto는 스파스 어텐션(sparse attention) + MoE를 활용하여 Orin X에서 10Hz 및 100ms의 레이턴시를 실현하고 있습니다. Xpeng의 2세대 VLA는 80ms 미만의 레이턴시를 실현하고 있습니다. DeepRoute.ai의 40B 파라미터 모델은 KV 캐시, 멀티 토큰 예측(MTP), 양자화 및 맞춤형 엔진을 활용하여 싱글 스텝 지연 시간을 60-85ms, 폐쇄 루프를 10-15Hz로 억제하고 있습니다.
연산 능력의 적응 : DeepRoute.ai는 100TOPS 플랫폼에 순수한 주행용 VA 모델을, 500TOPS 플랫폼에 추론 가능한 VLA 모델을 배포할 수 있습니다. Qualcomm 8797 칩 2개를 탑재한 Leapmotor D19(1280TOPS)는 단말 측에서 VLA 지원 주행을 실현하고 있습니다. Geely H9는 NVIDIA Thor 칩 2개를 채택하고 있습니다(2000TOPS).
오픈 루프 성능 : nuScenes 데이터셋을 기반으로 한 평가에서 VLA는 월드 모델에 비해 궤적 오차가 현저히 작은 것으로 나타났습니다. 예를 들어, 70억 개의 파라미터를 가진 AutoDrive-R2는 L2 거리 0.19m, SENNA는 0.22m를 달성했지만, 최고 성능의 월드 모델인 Drive-OccWorld조차도 0.32m에 그쳤습니다. 이러한 결과는 VLA가 인간의 실제 주행 궤적을 재현할 수 있는 능력을 갖추고 있음을 보여줍니다.
주요 과제
실시간 성능 및 연산 능력의 병목 현상 : 기존의 자기 회귀형 VLA 생성은 3-6Hz에 불과하며, 단일 추론의 지연 시간은 일반적으로 200ms를 초과하여 차량의 연산 능력과 스토리지 대역폭을 대량으로 소모합니다.
데이터 감독의 부족 : VLA는 고차원 시각 입력을 수신하는 반면, 저차원의 희소 액션에 의해 감독되기 때문에 모델의 잠재력이 제한됩니다. 고밀도 감독을 수행하기 위해서는 미래 이미지를 예측하는 월드 모델이 필요하거나, 동영상 예측 사전 학습을 채택해야 합니다.
안전성 중복성의 부족 : 엔드투엔드 단일 모델을 사용하는 VLA는 내결함성이 낮아, 기존 알고리즘을 통한 폴백이 필요합니다. 예를 들어, 널리 채택되고 있는 패스트-슬로우 듀얼 시스템, Horizon Robotics의 Lite Safety Checker, 그리고 Bosch의 세이프티 게팅 보상 메커니즘 등이 있습니다.
환각 및 롱테일 시나리오 : 복잡하거나 이전에 본 적 없는 시나리오에서는 성공률이 떨어집니다. 개선을 위해서는 월드 모델과 강화 학습을 통한 탐색이 필요하며, 예로 Bosch의 ExploreVLA, Huawei의 WEWA, Momenta의 R7 등을 들 수 있습니다.
OEM 각사 중에서는 Xpeng의 2세대 VLA와 Li Auto의 MindVLA가 대표적인 사례로 꼽히며, 이들은 각각 ‘네이티브 멀티모달 실세계 기반 모델’과 ‘공간·언어·행동의 통합 + 암묵적 세계 모델’이라는 두 가지 기술적 접근 방식을 채택하고 있습니다.
Xpeng의 2세대 VLA
Xpeng은 자율주행이 본질적으로 물리적인 AI 문제라는 견해를 내세우고 있습니다. 이 회사의 2세대 VLA는 네이티브 멀티모달 물리적 세계 기반 모델로 구축되었습니다. 네이티브 멀티모달 토큰라이저를 통해 단일 모달리티의 편향을 피하기 위한 고효율 초기 단계 융합을 실현하고 있습니다. 시각 추론의 ‘사고 연쇄(CoT)’를 통해 추론 효율이 32배 향상됩니다. 차간 추종 시나리오에서는 모델이 차선 변경이나 차간 추종과 같은 조작 방안을 자동으로 생성하고, 평가를 위한 추상적인 조감도를 작성합니다. 네이티브 콕핏·운전 연동을 통해 모델은 액션뿐만 아니라 동영상과 음성도 생성할 수 있으며, VLA의 기반이자 세계 모델, 시뮬레이션, 강화 학습을 위한 기반 프레임워크로 기능합니다. 2단계 모델을 참고하여 1단계 기반 모델을 재학습시키고, 언어 변환 링크를 제거함으로써 지연 시간을 80ms 미만으로 억제하고 있습니다.
Li Auto MindVLA
이 아키텍처는 3가지 구성 요소로 이루어져 있습니다.
V(공간 지능) : BEV 및 OCC를 기반으로, 중간 표현으로 3D 가우스 분포를 채택하고, LiDAR 포인트 클라우드를 3D 기하학적 프롬프트로 활용하여, 3D ViT 인코더와 피드포워드 3DGS를 통해 3D 씬 재구성을 수행합니다. 정적 환경과 동적 객체는 별도로 모델링됩니다.
L(언어 지능) : LLM 기반 모델을 재학습합니다(MoE 및 스파스 어텐션 메커니즘 활용). 고속 사고의 병렬 디코딩에서는 액션 토큰을 직접 출력하고, 저속 사고에서는 CoT와 액션 토큰을 동시에 출력합니다.
A(액션 전략) : 액션 전문가를 통합한 VLA-MoE 아키텍처를 채택했습니다. 이산 확산과 병렬 디코딩의 반복적 최적화를 통해 고정밀 주행 궤적을 출력합니다.
4단계 엔지니어링:
VL 기반 모델의 사전 학습(32B, Orin-X/Thor-U 듀얼 환경에 적응시키기 위해 3.6B의 MoE로 포스트 디스틸)
모방 학습을 통한 사후 훈련(4B)
강화 학습(RLHF + 순수 RL, 폐루프형 월드 시뮬레이터 구축)
드라이버 에이전트 HMI
그 후, MindVLA-U1을 통해 언어와 연속 동작의 통합 스트리밍 및 공동 모델링이 가능해졌으며, Intent-CFG가 도입되었습니다.
기타 OEM :
샤오미(Xiaomi) XLA 인지 라지 모델(VLM+엣지-클라우드 통합, VLA 도입 예정);
Leapmotor LEAP 4.0 VLA(엣지 측에서의 풀 모달리티, ‘이해·계획·미리보기·판단·수정’의 폐쇄 루프);
Great Wall CP Master Dual VLA(좌뇌 : 운전 에이전트, 우뇌 : 콕핏 에이전트);
Chery Falcon 900(VLA+월드 모델, L3 지원).
공급업체의 가장 대표적인 솔루션은 NVIDIA의 Alpamayo와 DeepRoute.ai의 40B VLA로, 각각 ‘추론 기반 듀얼 LLM VLA + 월드 모델’ 및 ‘40B 통합 기반 3단계 VLA’ 솔루션을 구현하고 있습니다.
NVIDIA Alpamayo
NVIDIA Alpamayo는 물리 AI 데이터셋, VLA 대규모 모델, 그리고 AlpaSim 시뮬레이션 프레임워크를 아우르는 시스템 수준의 VLA 솔루션입니다. 각 OEM사에 차량 도입을 위한 3세대의 궤적 생성 옵션을 제공합니다. 구체적으로는 회귀 예측(TensorRT, 예 : SparseDrive), 확산 생성(TensorRT, 예 : DiffusionDrive), 그리고 플로우 매칭(TensorRT-Edge-LLM, 예 : Qwen3 VL 기반 Alpamayo)이 있습니다.
Alpamayo-R1-10B를 예로 들어보겠습니다. 입력에는 과거 이미지, 사용자 지시, 과거 궤적 및 노이즈가 포함된 액션이 포함됩니다. 입력 데이터는 VLM 플로우 매칭 토큰라이저를 통해 텍스트, 이미지 및 궤적 토큰으로 변환됩니다. 80억 개의 파라미터를 가진 Qwen3 VL-LLM이 장면 이해와 암묵적 CoT 추론을 수행하여 추론 텍스트를 출력함과 동시에 KV 캐시를 생성합니다. 2B 매개변수를 가진 Qwen3 VL-LLM은 노이즈가 섞인 궤적과 KV 캐시를 받아, 플로우 매칭을 통한 단계적인 노이즈 제거 및 보정을 수행한 후, 최종적으로 미래의 궤적을 출력합니다. Cloud Cosmos 월드 모델은 롱테일 시나리오를 위한 훈련 데이터를 생성합니다. 차량에 도입할 때는 안전상의 폴백 수단으로 기존 알고리즘을 채택하고, 엔드투엔드 VLA를 주요 시스템으로 하여 L2-L4를 지원합니다. 향후 잠재 공간 추론의 처리 속도는 2-4배 향상될 것으로 기대됩니다.
DeepRoute.ai 40B VLA
DeepRoute.ai는 자율주행의 의사결정을 3단계로 구분합니다.
관찰 단계-멀티 카메라 영상이 약 1,000개의 시각 토큰으로 인코딩됩니다.
추론 단계-모델은 시나리오에 대한 상세한 의미 분석을 수행하여 주요 이벤트 및 의사결정 로직에 대한 설명을 생성합니다. 추론 토큰의 수는 10-50개로 엄격하게 제어됩니다.
실행 단계-주행 제어 명령을 출력하는 데 필요한 토큰은 불과 10개 정도입니다.
본 시스템은 통합된 40B 매개변수의 기반 모델을 통해 ‘운전자(센서 입력을 바탕으로 행동함)’, ‘분석가(인과 관계를 분석함)’, ‘해설자(판단하고 의사결정을 내림)’라는 3가지 기능을 통합하고 있습니다. ‘운전 전 사고’를 실현하기 위해 V+A, V+A→L, V→L+A라는 3가지 작업 범주에 걸친 공동 학습이 실시되고 있습니다. 사전 학습에서는 궤적에 기반한 지도 학습에서 동영상 예측으로 전환합니다. 방대한 양의 동영상을 활용하여 픽셀 단위로 물리 법칙을 학습함으로써, 데이터 활용률을 0.001%에서 100%로 향상시킵니다. 배포 시에는 KV 캐시, MTP, 양자화 및 맞춤형 추론 엔진을 통해 단계당 지연 시간을 60-85ms로 줄여, 10-15Hz의 차량 실시간 폐루프를 실현합니다. 모델 증류는 연산 능력에 따라 수행되며, 순수한 주행용 VA 모델은 100TOPS 플랫폼에서 완전한 VLA 모델은 500TOPS 플랫폼에서 작동합니다.
기타 공급업체:
Afari Technology는 ‘VLA+E2E’ 협업형 폐루프를 채택하고 있습니다. VLA의 저속 시스템은 CoT 텍스트를 출력하고, E2E의 고속 시스템은 표적 감지/차선 감지 결과를 출력하여, 이를 융합해 계획 및 제어에 활용합니다. StepVL 4.0은 32B 사전 학습 모델에서 7B로 증류되었습니다.
QCraft는 2026년에 ‘VLA+월드 모델+강화 학습’이라는 통합 아키텍처로 업그레이드하여, 단일 Journey 6M 칩 상에서 도시 지역의 NOA를 구현할 예정입니다.
Zhuoyu는 VLA 월드 모델(네이티브 멀티모달 기반 모델, Chain of World, 그리고 구조와 운동이 분리된 잠재 운동 표현)을 출시했습니다.
동향 1 : 하이브리드 아키텍처가 주류로 부상
디퓨전, 트랜스포머, LLM/VLM의 심층적인 통합이 주류가 될 것입니다. 디퓨전은 고품질의 연속적인 동작 및 궤적 생성에 탁월합니다. 트랜스포머는 긴 시퀀싱 모델링에 뛰어나며, LLM/VLM은 의미론적 및 멀티모달 이해에 탁월합니다. 대표적인 예로는 Li Auto의 MindVLA(3D 가우스 분포 + MoE LLM + 확산·액션·전문가), NVIDIA의 GR00T-N1(패스트·슬로우 듀얼 시스템 : 패스트 200Hz 확산·액션,슬로우 10Hz VLM), HybridVLA(자기회귀 + 협업형 확산) 등이 있습니다. 업계에서는 다음과 같은 3가지 통합 모델이 형성되어 있습니다. 1. 단일 모델의 엔드투엔드(E2E) + 월드 모델 + 강화 학습(RL)(Momenta, Horizon Robotics); 2. VLA + 월드 모델(XPeng 등); 3. E2E+VLM/VLA 기반 모델(Afari Technology의 VLA 슬로우 시스템+E2E 패스트 시스템).
동향 2 : VLA와 범용 월드 모델의 통합을 통한 ‘월드 VLA/VLA 월드 모델’의 탄생
VLA는 인식과 행동을 담당하고, 월드 모델은 예측을 담당합니다. 이를 통합함으로써 자동차는 단순한 이동 수단에서 이동 로봇으로 변모하며, 규칙 주도형에서 인식 주도형으로 전환되어 자율적인 지각, 추론·의사결정, 그리고 정밀한 실행 능력을 갖추게 됩니다. Zhuoyu의 VLA 월드 모델은 3세대 네이티브 멀티모달 기반 모델로 진화했습니다. ‘Chain of World’를 통해 잠재 공간에서 다단계 세계 상태 예측을 수행함으로써 ‘행동 전에 사고한다’는 것을 실현하고 있습니다. 구조와 운동의 분리, 그리고 잠재적 운동 표현을 통해 재구성 비용을 절감하고 있습니다. Geely G-ASD는 VLA와 월드 모델을 통합하여 차량이 자동으로 작업을 수행할 수 있도록 합니다. WorldVLA는 생성을 위해 동작과 이미지를 공동으로 이해하며, 월드 모델과 동작 모델이 서로를 보완합니다.
동향 3 : VLA + 월드 모델 + 강화 학습(RL)의 삼위일체 통합, RL을 핵심 엔진으로
업계에서는 ‘사전 학습 → 시뮬레이션 → 강화 학습’이라는 3층 아키텍처가 형성되고 있습니다. 월드 모델이 롱테일 시나리오를 생성하고, VLA가 인루프 추론을 수행하며, 강화 학습이 추론 공간에서 최적의 전략을 반복적으로 도출합니다. 대표적인 예로는 화웨이의 WEWA 2.0(다중 에이전트 게임 + 클라우드 온라인 RL, 학습 강도가 10배 증가), Momenta의 R7(3단계 프로세스 : 사전 학습 → 시뮬레이션 → RL, AI를 ‘모방자’에서 ‘의사결정자’로 변혁),Pony.ai의 PonyWorld 2.0(자가 진단 + 목표 지향적 진화 + 정밀 플라이휠) 등이 있습니다.
한편, 월드 모델은 픽셀 수준의 예측에서 잠재 공간 및 인과 추론으로 진화하고 있습니다. NVIDIA의 Alpamayo는 잠재 공간에서의 암묵적 추론을 통해 2-4배의 속도 향상을 실현하며, Chain of Causality(CoC)를 통해 완전한 추론 체인을 생성합니다. Li Auto는 예측적 암묵적 월드 모델을 VLA에 통합했습니다. Xpeng은 언어 변환 단계를 제거하고 정보 손실을 줄이기 위해 아키텍처를 V-L-A에서 V/L-A로 개정했습니다. Huawei의 DriveVLA-W0는 월드 모델 통합을 통해 데이터 양이 70만 프레임에서 7,000만 프레임으로 확대됨에 따라 충돌률이 지속적으로 감소하며, 이러한 이점이 증폭되어 데이터 스케일링 법칙이 강화됨을 입증하고 있습니다.
동향 4 : 기술 구현 및 안전성 확보 가속화
2025년부터 2026년에 걸쳐, 단일 모델 엔드투엔드 및 VLA 솔루션이 널리 도입될 것입니다. L3/L4의 진화에 따라 안전성 중복성이 필수적입니다(기존 알고리즘을 통한 폴백 + 엔드투엔드 메인 시스템. 예 : NVIDIA의 패스트-슬로우 듀얼 시스템, Horizon Robotics의 Lite Safety Checker, Bosch의 안전 게이팅 PDMS 보상). 연산 능력의 계층적 디스티레이션이 양산화의 핵심이 되고 있습니다. 100-500 TOPS 수준의 VA/VLA(DeepRoute.ai)의 유연한 전개, 그리고 L3를 지원하는 듀얼 Thor/듀얼 8797 등의 고성능 컴퓨팅 플랫폼입니다.
VLA는 자율주행이 ‘엔드투엔드 지각·제어’에서 ‘이해·추론·제어’로 진화하기 위한 핵심적인 경로가 될 것입니다. 2026년에는 OEM(XPeng, Li Auto 등)과 공급업체(NVIDIA, DeepRoute.ai, Afari Technology, QCraft 등)의 공동 추진을 통해 VLA는 월드 모델 및 강화 학습과 깊이 통합되어 하이브리드 아키텍처를 형성하고, ‘World VLA’의 프로토타입이 구체화될 것입니다. 지연 시간, 연산 능력, 데이터 감독, 보안 중복성 등의 과제가 점차 해결됨에 따라, VLA는 레벨 3 이상의 자율주행의 대규모 확장을 뒷받침하며, 차량이 물리적 세계에서의 범용 에이전트로 진화할 수 있게할 것입니다.
Research on Automotive and Robot VLA: Hybrid Architectures Become Mainstream, VLA Integrates with General World Models, and Reinforcement Learning Serves as Core Engine
Vision-Language-Action (VLA) model is a model integrating vision, language and action modalities. Adopting a unified multimodal learning framework, it integrates perception, reasoning and control, and generates executable physical world actions (e.g., robot joint motion, and vehicle steering/acceleration/braking control) directly from visual inputs (images/videos) and language instructions.
VLA equals VLM (for understanding and description) plus E2E (end-to-end decision and control) plus CoT (chain-of-thought human-like reasoning). In terms of capability comparison, VLA delivers precise 3D perception, commonsense understanding, logical thinking and interpretability simultaneously.
The evolution of VLA in autonomous driving falls into four stages.
Language as interpreter (Pre-VLA): Language models only generate scene descriptions without participating in control.
Modular VLA: Language acts as a planning component for decision, yet multi-stage workflows incur latency.
Unified end-to-end VLA: Sensor inputs are directly mapped to actions via a single forward propagation.
Reasoning-enhanced VLA: LLMs enter the control closed loop to enable long-term reasoning, memory and interaction capabilities, exemplified by Li Auto MindVLA.
Implementation timeline:
Segmented end-to-end came into mass production from 2024 to 2025.
One-model end-to-end and VLA were largely rolled out between 2025 and 2026.
The year 2026 marks a critical window period of "intensified multi-route competition and accelerated paradigm integration" for intelligent driving large models.
Technical status and core indicators.
Wide-ranging model parameters: NVIDIA Alpamayo 1.5 includes 0.5B/10B parameters; DeepRoute.ai uses a 40B-parameter foundation model; StepVL foundation model from Afari Technology features 32B parameters (distilled to 7B and 3.6B); Li Auto MindVLA 32B-parameter foundation model is distilled into a 3.6B-parameter MoE variant, 4B for vehicle model.
In-vehicle real-time performance: Li Auto leverages sparse attention + MoE to realize 10Hz and 100ms latency on Orin X; Xpeng's second-generation VLA enables <80ms latency; DeepRoute.ai's 40B-parameter model uses KV Cache, Multi-Token Prediction (MTP), quantization and customized engine to achieve single-step latency of 60-85ms and 10-15Hz closed loop.
Computing power adaptation: DeepRoute.ai can deploy pure driving VA models on 100TOPS platforms and reasoning-capable VLA models on 500TOPS platforms; Leapmotor D19 equipped with dual Qualcomm 8797 chips (1280TOPS) realizes end-side VLA-assisted driving; Geely H9 adopts dual NVIDIA Thor chips (2000TOPS).
Open-loop performance: Based on the nuScenes dataset, VLA exhibits notably lower trajectory errors than world models. For instance, AutoDrive-R2 with 7B parameters delivers an L2 distance of 0.19m and SENNA achieves 0.22m; the best-performing world model Drive-OccWorld reaches 0.32m. Such results demonstrate VLA's ability to reproduce human real-world driving trajectories.
Major challenges
Real-time performance and computing power bottlenecks: Traditional auto-regressive VLA generation only reaches 3-6Hz, and single-reasoning latency commonly exceeds 200ms, consuming a lot of vehicle computing power and storage bandwidth.
Data supervision deficit: VLA receives high-dimensional visual inputs yet is supervised by low-dimensional sparse actions, limiting model potential. World models are required to predict future images for dense supervision, or video prediction pre-training shall be adopted.
Lack of safety redundancy: End-to-end single-model VLA has low fault tolerance and requires fallback from traditional algorithms, e.g., the widely adopted fast-slow dual-system, Horizon Robotics Lite Safety Checker and Bosch safety gating reward mechanism.
Hallucinations and long-tail scenarios: Success rates drop in complex or unseen scenarios. World model + reinforcement learning exploration is required to improve, e.g., Bosch ExploreVLA, Huawei WEWA and Momenta R7.
Among OEMs, Xpeng's second-generation VLA and Li Auto MindVLA stand as representative cases, corresponding respectively to two technical routes, namely "native multimodal physical world foundation model" and "space-language-action unification + implicit world model".
Xpeng's Second-Generation VLA
Xpeng holds to the view that intelligent driving is essentially a physical AI problem. Its second-generation VLA is built as a native multimodal physical world foundation model. A native multimodal tokenizer enables highly efficient early-stage fusion to avoid single-modality bias. Visual reasoning chain-of-thought (CoT) boosts reasoning efficiency by 32 times. In car following scenarios, the model automatically generates maneuver proposals such as lane change or car following, and produces abstract bird-eye-view diagrams for scoring. Native cockpit-driving linkage allows the model to generate not only actions but also videos and sounds, serving as the foundation for VLA and the foundation framework for world models, simulation, and reinforcement learning. Referring to two-stage models, it retrains one-stage foundation models and eliminates language translation links to bring latency below 80ms.
Li Auto MindVLA
Its architecture consists of three components.
V (spatial intelligence): Based on BEV and OCC, it adopts 3D Gaussian as intermediate representation, leverages LiDAR point clouds as 3D geometric prompts, and performs 3D scene reconstruction via 3D ViT encoder and feed forward 3DGS. Static environments and dynamic objects are modeled separately.
L (language intelligence): Retrains LLM foundation model (leveraging MoE and sparse attention mechanism). Fast-thinking parallel decoding directly outputs Action Tokens, while slow thinking outputs CoT and Action Tokens simultaneously.
A (action strategy): Adopts VLA-MoE architecture embedded with Action Expert. Discrete diffusion and parallel decoding iterative optimization are used to output high-precision driving trajectories.
Four-phase engineering:
VL foundation model pre-training (32B, post-distilled to 3.6B MoE to adapt to dual Orin-X/Thor-U)
Imitation learning post-training (4B)
Reinforcement training (RLHF + pure RL, building a closed-loop world simulator)
Driver agent HMI
Subsequent MindVLA-U1 enables unified streaming and joint modeling of language and continuous actions, and introduces Intent-CFG.
Other OEMs:
Xiaomi XLA Cognitive Large Model (VLM + edge-cloud integration, VLA planned);
Leapmotor LEAP 4.0 VLA (edge-side full-modality, "understanding-planning-preview-judgment-correction" closed loop);
Great Wall CP Master Dual VLA (left brain driving agent, right brain cockpit agent);
Chery Falcon 900 (VLA + world model, supporting L3).
The most typical solutions of suppliers are NVIDIA Alpamayo and DeepRoute.ai 40B VLA, representing "reasoning-based dual-LLM VLA + world model" and "40B unified base three-stage VLA" solution respectively.
NVIDIA Alpamayo
NVIDIA Alpamayo is a system-level VLA solution encompassing physical AI dataset, VLA large model and AlpaSim simulation framework. It provides OEMs with three generations of trajectory generation options for vehicle deployment: regression prediction (TensorRT, e.g., SparseDrive), diffusion generation (TensorRT, e.g., DiffusionDrive), and flow matching (TensorRT-Edge-LLM, e.g., Alpamayo based on Qwen3 VL).
Take Alpamayo-R1-10B as an example. Inputs include historical images, user instructions, historical trajectories and noisy actions. Inputs are converted into Text, Image and Trajectory Tokens via VLM Flow Matching Tokenizer. The 8B-parameter Qwen3 VL-LLM handles scene understanding and implicit CoT reasoning, outputs reasoning texts and generates KV Cache. The 2B-parameter Qwen3 VL-LLM receives noisy trajectories and KV Cache, conducts progressive denoising and correction via Flow Matching, and finally outputs future trajectories. Cloud Cosmos world model generates training data for long-tail scenarios. Vehicle deployment adopts traditional algorithm as safety fallback + end-to-end VLA as primary system, supporting L2-L4. Future latent space reasoning is expected to accelerate by 2-4 times.
DeepRoute.ai 40B VLA
DeepRoute.ai breaks down autonomous driving decision into three phases.
Observation phase - Multi-camera videos are encoded into approximately 1,000 visual tokens.
Reasoning phase - The model conducts in-depth semantic analysis of scenarios and generates descriptions of key events and decision logics, with the number of reasoning tokens strictly controlled within 10-50.
Execution phase - Outputting driving control commands requires only about 10 tokens.
It integrates three capabilities of "driver (acting based on sensor inputs), analyst (analyzing causality), and commentator (judging and making decisions)" through a unified 40B-parameter foundation model. Joint training is implemented across three task categories, namely, V+A, V+A->L and V->L+A, to realize "thinking before driving". Pre-training switches from trajectory supervision to video prediction. Massive videos are leveraged to learn physical laws at per-pixel level, lifting data utilization rate from 0.001% to 100%. During deployment, KV Cache, MTP, quantization and customized reasoning engines reduce single-step latency to 60-85ms to achieve a 10-15Hz vehicle real-time closed loop. Model distillation is performed according to computing power: pure driving VA models run on the 100TOPS platform, and complete VLA models operate on the 500TOPS platform.
Other suppliers:
Afari Technology adopts the "VLA+E2E" collaborative closed loop. The VLA slow system outputs CoT texts while the E2E fast system outputs target detection/lane detection results for fusion into planning and control. StepVL 4.0 is distilled from 32B pre-trained to 7B.
QCraft upgrades to the "VLA + world model + reinforcement learning" unified architecture in 2026, and realizes urban NOA on single Journey 6M chip.
Zhuoyu launches VLA World Model (native multimodal foundation model, Chain of World, and structure-motion decoupled latent motion representation).
Trend 1: Hybrid Architectures Become Mainstream
Deep integration of Diffusion, Transformer and LLM/VLM becomes mainstream. Diffusion excels at generating high-quality continuous actions and trajectories. Transformer is good at long sequence modeling. LLM/VLM is skilled in semantic and multimodal understanding. Representative examples include Li Auto MindVLA (3D Gaussian + MoE LLM + Diffusion Action Expert), NVIDIA GR00T-N1 (Fast-Slow Dual System: Fast 200Hz Diffusion Action, Slow 10Hz VLM), HybridVLA (Autoregression + Collaborative Diffusion). The industry has formed three integration models: 1. One-model end-to-end + world model + RL (Momenta, Horizon Robotics); 2. VLA + world model (XPeng, etc.); 3. E2E + VLM/VLA foundation model (Afari Technology VLA slow system + E2E fast system).
Trend 2: VLA and General World Model Integrate into World VLA / VLA World Model
VLA undertakes cognition and action while the world model takes on future prediction. Their unification transforms automobiles from transportation means into mobile robots, shifting from rule-driven to cognition-driven, with capabilities of autonomous perception, reasoning & decision and precise execution. Zhuoyu's VLA World Model has evolved into its third-generation native multimodal foundation model. With Chain of World, it performs multistep world state prediction in latent space, achieving "thinking before acting." Structure-motion decoupling and latent motion representation reduce reconstruction costs. Geely G-ASD integrates VLA and world model, enabling vehicles to automatically perform tasks. WorldVLA jointly understands actions and images for generation, with the world model and action model mutually reinforcing each other.
Trend 3: VLA + World Model + Reinforcement Learning (RL) Trinity Integration, with RL as the Core Engine
The industry forms a "pre-training -> simulation -> reinforcement learning" three-layer architecture. The world model generates long-tail scenarios, VLA conducts in-loop reasoning, and reinforcement learning iterates optimal strategies in the inference space. Representative examples include: Huawei WEWA 2.0 (multi-agent gaming + cloud online RL, training intensity increased by 10 times); Momenta R7 (three-stage process: pre-training -> simulation -> RL, turning AI from "imitator" to "decision-maker"); Pony.ai's PonyWorld 2.0 (self-diagnosis + targeted evolution + precision flywheel).
Meanwhile, world models evolve from pixel-level prediction toward latent space and causal reasoning. NVIDIA Alpamayo achieves 2-4-fold acceleration via implicit reasoning in the latent space, and generates a complete reasoning chain through Chain of Causality (CoC). Li Auto embeds predictive implicit world models into VLA. Xpeng eliminates language translation links and revises architecture from V-L-A to V/L-A to mitigate information loss. Huawei DriveVLA-W0 verifies that with world model integration, collision rates keep decreasing as data volume expands from 0.7 million to 70 million frames and such advantages are amplified, strengthening the data scaling law.
Trend 4: Engineering Implementation and Safety Assurance Accelerate
One-model end-to-end and VLA solutions are largely implemented from 2025 to 2026. The evolution of L3/L4 has driven safety redundancy to become a necessity (traditional algorithm fallback + end-to-end main system, e.g., NVIDIA's fast-slow dual systems, Horizon Robotics' Lite Safety Checker, and Bosch's safety gating PDMS reward). Hierarchical distillation of computing power has become key to mass production: flexible deployment of VA/VLA (DeepRoute.ai) at 100-500 TOPS, and high-performance computing platforms such as dual Thor/dual 8797 supporting L3.
VLA serves as the core route for intelligent driving to evolve from "end-to-end perception-control" toward "understanding-reasoning-control". In 2026, driven by both OEMs (XPeng, Li Auto, etc.) and suppliers (NVIDIA, DeepRoute.ai, Afari Technology, QCraft, etc.), VLA is deeply integrated with world model and reinforcement learning, forming a hybrid architecture, and the prototype of World VLA takes shape. As latency, computing power, data supervision, and security redundancy issues are gradually resolved, VLA will support the large-scale deployment of L3 and above autonomous driving and enable vehicles to evolve into general agents in the physical world.
Definitions