Site Search

In recent years, advancements in AI have expanded beyond text and image processing to include robot control. Among these advancements, the Vision-Language-Action (VLA) model is attracting particular attention.

VLA is a new AI architecture that understands visual information (Vision) acquired from a camera and instructions in natural language (Language), and directly converts them into robot actions (Action). In other words, it's a system that connects the robot's "seeing," "understanding," and "moving" into a single model.

This articleintroduces NVIDIA platforms suitable for Physical AI training and robot inference, based on actual measurement results of the "NVIDIA Isaac GR00T N1.7".

What is NVIDIA Isaac GR00T N1.7?

NVIDIA's Isaac GR00T N1.7 is an open VLA model designed to enable general-purpose humanoid robot skills. It receives natural language instructions and camera images as input and generates robot movements. It also supports fine-tuning with proprietary data, allowing for flexible adaptation to specific robots, tasks, and environments.


Reference: Isaac-GR00T N1.7 Github

    NVIDIA Isaac GR00T

Details of this verification

When using GR00T N1.7, the following points are important when selecting a development environment: "Which GPU environment will be used for training and evaluation?" and "To which platform will the completed model be deployed?"

Therefore, in this study, we fine-tuned the GR00T N1.7-3B on three​ ​NVIDIA platforms using the same dataset and training conditions. We compared the training times in each environment to determine which platform to choose based on your application.

Compared environments

platform

Main uses

composition

NVIDIA® Jetson Thor™

Real-time inference on robots,

Sensor processing, control

14-core CPU

128GB memory

NVIDIA DGX Spark™

Development of a prototype generative AI model,

Inference, fine tuning

20 CPU cores

128GB Unified Memory

NVIDIA RTX PRO™ 6000 Blackwell Workstation Edition

AI model training and evaluation,

simulation

64-core CPU

GPU memory 96GB

For more information on NVIDIA GPUs and related hardware, please click here.

Fi tuning conditions

Base model

GR00T N1.7-3B

dataset

DROID Dataset(lerobot/droid_1.0.1)

Usage data

1,000 Episodes

Comparison items

Time to complete 10,000 Steps fine tuning

inspection result

In this test, the RTX PRO 6000 Blackwell was the fastest. This configuration is recommended for physical AI development involving iterative training and evaluation.

platform

Study time

Jetson Thor

Approximately 12 hours

DGX Spark

Approximately 15 hours

RTX PRO 6000 Blackwell

Approximately 2.5 hours

The RTX PRO 6000 completes training in approximately 5 to 6 times less time compared to Jetson Thor and DGX Spark.

Why is study time important?

Robot development isn't a one-time process of fine-tuning. We continuously improve the model by evaluating checkpoints generated during the learning process and adjusting data and parameters as needed.

In development that involves repeatedly going through this cycle, training time is not simply waiting time. Longer training times increase the interval between checkpoint evaluations and the next training cycle, reducing the number of trials that can be performed within a limited development period. This, in turn, affects the speed of model improvement. In short, the length of training time is a crucial factor influencing the overall speed of the development cycle.

By reducing learning time, we can iterate more frequently on evaluating and improving Checkpoints, thereby accelerating the entire development cycle.

Summary

This article introduces the VLA model, a technology for realizing Physical AI, and provides an overview of the NVIDIA Isaac GR00T N1.7, while also comparing the fine-tuning times on Jetson Thor, DGX Spark, and RTX PRO 6000. The measured results demonstrate that choosing the appropriate platform for each application is key to efficient Physical AI development.
*For large datasets or when using multiple GPUs, higher-end configurations are also an option.

We hope this article will be helpful when considering the training time and physical AI development environment for GR00T N1.7.

Related page

For applications of DGX Spark for generative AI, please see the following article.

Contact Us

If you are considering implementing Physical AI, or if you have any questions regarding building or deploying a learning environment, please feel free to contact us.