In recent years, advancements in AI have expanded beyond text and image processing to include robot control. Among these advancements, the Vision-Language-Action (VLA) model is attracting particular attention.
VLA is a new AI architecture that understands visual information (Vision) acquired from a camera and instructions in natural language (Language), and directly converts them into robot actions (Action). In other words, it's a system that connects the robot's "seeing," "understanding," and "moving" into a single model.
This articleintroduces NVIDIA platforms suitable for Physical AI training and robot inference, based on actual measurement results of the "NVIDIA Isaac GR00T N1.7".
What is NVIDIA Isaac GR00T N1.7?
NVIDIA's Isaac GR00T N1.7 is an open VLA model designed to enable general-purpose humanoid robot skills. It receives natural language instructions and camera images as input and generates robot movements. It also supports fine-tuning with proprietary data, allowing for flexible adaptation to specific robots, tasks, and environments.
Reference: Isaac-GR00T N1.7 Github
Details of this verification
When using GR00T N1.7, the following points are important when selecting a development environment: "Which GPU environment will be used for training and evaluation?" and "To which platform will the completed model be deployed?"
Therefore, in this study, we fine-tuned the GR00T N1.7-3B on three NVIDIA platforms using the same dataset and training conditions. We compared the training times in each environment to determine which platform to choose based on your application.
Compared environments
|
platform |
Main uses |
composition |
|
NVIDIA® Jetson Thor™ |
Real-time inference on robots, Sensor processing, control |
14-core CPU 128GB memory |
|
NVIDIA DGX Spark™ |
Development of a prototype generative AI model, Inference, fine tuning |
20 CPU cores 128GB Unified Memory |
|
NVIDIA RTX PRO™ 6000 Blackwell Workstation Edition |
AI model training and evaluation, simulation |
64-core CPU GPU memory 96GB |
For more information on NVIDIA GPUs and related hardware, please click here.
Fi tuning conditions
|
Base model |
GR00T N1.7-3B |
|
dataset |
DROID Dataset(lerobot/droid_1.0.1) |
|
Usage data |
1,000 Episodes |
|
Comparison items |
Time to complete 10,000 Steps fine tuning |
inspection result
In this test, the RTX PRO 6000 Blackwell was the fastest. This configuration is recommended for physical AI development involving iterative training and evaluation.
|
platform |
Study time |
|
Jetson Thor |
Approximately 12 hours |
|
DGX Spark |
Approximately 15 hours |
|
RTX PRO 6000 Blackwell |
Approximately 2.5 hours |
The RTX PRO 6000 completes training in approximately 5 to 6 times less time compared to Jetson Thor and DGX Spark.
Why is study time important?
Robot development isn't a one-time process of fine-tuning. We continuously improve the model by evaluating checkpoints generated during the learning process and adjusting data and parameters as needed.
In development that involves repeatedly going through this cycle, training time is not simply waiting time. Longer training times increase the interval between checkpoint evaluations and the next training cycle, reducing the number of trials that can be performed within a limited development period. This, in turn, affects the speed of model improvement. In short, the length of training time is a crucial factor influencing the overall speed of the development cycle.
By reducing learning time, we can iterate more frequently on evaluating and improving Checkpoints, thereby accelerating the entire development cycle.
Summary
This article introduces the VLA model, a technology for realizing Physical AI, and provides an overview of the NVIDIA Isaac GR00T N1.7, while also comparing the fine-tuning times on Jetson Thor, DGX Spark, and RTX PRO 6000. The measured results demonstrate that choosing the appropriate platform for each application is key to efficient Physical AI development.
*For large datasets or when using multiple GPUs, higher-end configurations are also an option.
We hope this article will be helpful when considering the training time and physical AI development environment for GR00T N1.7.
Related page
For applications of DGX Spark for generative AI, please see the following article.
Contact Us
If you are considering implementing Physical AI, or if you have any questions regarding building or deploying a learning environment, please feel free to contact us.