What needs to change in the equipment used in an era when AI not only gives a single answer, but also calls tools and fixes code on its own?
The reason Nvidia’s Vera Rubin announcement is drawing attention is not simply that it involves faster chips.
Whether these figures will translate directly into lower costs for actual services must be examined separately.
3-Line Summary
1. Vera Rubin is a system for AI agents.
2. Nvidia presented 30x higher throughput per unit of power.
3. The initial performance figures have not yet undergone independent review.
A “bundle of tasks” has become more important than a single answer
Vera Rubin is Nvidia’s next-generation AI infrastructure. Here, an AI agent does not stop at answering a question with a single sentence; it carries out multiple steps, such as executing code, calling tools, and processing data. As tasks become longer, the system must continue to retain prior conversations and materials, making it difficult to judge perceived performance based on computational speed alone.
Nvidia’s AgentX measurement also targets this point. This benchmark by SemiAnalysis reproduces prerecorded coding sessions, testing long workflows rather than a single question and answer. Context accumulates, and requests may pause while tools are being executed. Results are also affected by the use of a KV cache, which prevents already processed content from being rewritten, and concurrent execution, which runs multiple tasks together.
This is why throughput per megawatt has moved ahead of the question of how many times faster a GPU has become. Because data centers operate within power and cooling infrastructure, how many useful tasks can be completed with the same amount of power may be a more direct metric for operators.
What do 30x and 1/35 mean?
According to reporting by Hankook Ilbo, Nvidia announced that Vera Rubin NVL72 delivers up to 30x higher throughput per unit of power than Blackwell GB300 NVL72 in actual agentic coding tasks. It also said that the token cost of generating 100만 tokens (1 million tokens) had been reduced to as low as 1/35.
However, it is difficult to interpret these figures as meaning that AI service fees will immediately become 1/35 of their current level. The announcement’s benchmarks are facility power and the operating cost of generating tokens, while service prices also include other costs such as software, data, personnel, and networks. The key point to note in this announcement is not the striking multiple itself, but that costs were presented in connection with task completion volume.
In addition, these results are initial measurements disclosed by Nvidia. Hankook Ilbo reported that external review is underway. Whether these are results of direct comparisons with other equipment in the same environment, and how much of this efficiency will continue in actual customer services, must be confirmed through subsequent materials. Initial benchmarks show potential, but do not by themselves establish a final advantage.
Orders and the space project are still “plans”
News surrounding Vera Rubin does not stop with performance announcements. According to reporting by Alpha Economy and NewsPim, Indian AI infrastructure company AM Intelligence said it had ordered 9,000 systems (9000 systems) and plans to begin operating servers in southern India starting next year. This is the company’s announcement and construction plan, not performance from a facility already in operation.
SpaceXAI has also outlined a plan to apply Vera Rubin NVL72 to AI satellites. According to AI Times, the first system is targeting orbital deployment in the fourth quarter of 2027, and joint design is underway to meet constraints involving power, thermal management, and communications bandwidth. This, too, should be viewed as a plan for space deployment.
The key to reading Vera Rubin is not simply that “a next-generation chip has arrived.” As AI takes on more complex work, the amount of work that can be completed within a given power budget is becoming a competitive benchmark. It is safer to treat the announced figures as a starting point and check whether actual operating cases and independent comparison results follow.
Criteria for reading the next announcement
It is difficult to reach a conclusion about Vera Rubin-related news based on a single number or sentence. The meaning becomes clearer when distinguishing who is making the announcement, whether the statement concerns an already implemented fact or a future plan, and whether the target and timing are specified concretely.
Even within the same material, explanations of need, discussions, implementation plans, and actual implementation may represent different stages. Even when the scale of an announcement appears large, it is necessary to check what it covers and what procedures remain in order to avoid overstating or understating the current situation.
In subsequent news, compare whether a new announcement repeats existing information or whether the target, schedule, or implementation status has actually changed. Reviewing the original sources listed in the references can also reveal differences in wording that are easy to miss from headlines alone.
References
Tags #VeraRubin #Nvidia #RubinGPU #AIAgents #AgenticAI #AISemiconductors #AIDataCenters #PowerEfficiency #AIInference #NVL72 #Blackwell #GPU #VeraCPU #AIInfrastructure