A certificate will not currently affect your device replacement schedule.Article author and source: GeekPark
Back in late July, if you were following the smartphone industry, you would have noticed that numerous smartphone manufacturers announced their devices had received National Artificial Intelligence Level 3 certification.
The starting point of this collective announcement was the release of the first batch of AI terminal intelligence grading test results on July 17, 2026.
According to the initial test list, 11 mobile devices have achieved L3, including 9 smartphones and 2 tablets. The smartphone lineup includes brands such as Huawei, Motorola, Honor, vivo, OPPO, Xiaomi, and Step星辰.
After the list was released, manufacturers quickly translated their technical test results into social media-friendly language. Huawei emphasized “first certification,” Xiaomi highlighted “national-level,” and vivo directly labeled L3 as the highest intelligent rating currently available. Huawei and Xiaomi officially announced their results on July 18, with OriginOS following the next day; similar L3 posters formed an unusual collective pre-launch campaign.
Although there are slight differences in wording, it is clear that these promotions all point to the same certification test: the "Classification of Artificial Intelligence Terminal Intelligence" released by the Ministry of Industry and Information Technology in May.
What exactly does this L3 capability mean? When will we be able to use these L3 phones? And if we interpret this according to automotive standards, does the longer-term L4 level represent a device form factor with fully autonomous AI agents requiring no human intervention?
L3 does not represent "autonomous driving"
Before introducing the entire certification standard, there are two common misunderstandings to avoid.
First, the certification corresponds to GB/Z 177-2026, "Classification of Intelligence for AI Terminals," which is a set of national standard guiding technical documents, not a mandatory national standard.
It provides a commonly used industry benchmark, not a strict准入 threshold determining whether a phone can be launched. This is also why the term “national standard” in the title is in quotation marks: while it does belong to the national standard system, it currently primarily serves a guiding and evaluative role.
Second, L3 evaluates a specific terminal device and its software and hardware versions, not the entire brand as a whole. A model passing L3 does not mean all phones from the same brand are automatically upgraded; not appearing on the initial list does not indicate test failure, as the first batch of results comes from manufacturers voluntarily submitting devices for testing, not random market sampling. This also clearly does not include all AI Agent phones on the market: for example, OPPO’s AI Agent phone project, Xiao Bu NEXT, which announced its pre-launch and entry into closed testing at the end of July, is not on this list.
In other words, L3 can be a product selling point, but it is not yet a comprehensive ranking in the mobile phone industry.
How to "upgrade" from L1 to L3
Additionally, you may have noticed that this standard is not actually a test of model capabilities. It breaks down terminal capabilities into five dimensions—perception, cognition, execution, memory, and learning—and further divides them into 14 specific abilities, such as user perception, task planning, tool invocation, long- and short-term memory, and contextual adaptation.
Putting these terms into a everyday scenario makes them easier to understand. Imagine the user telling their phone: “Help me plan a trip to Hangzhou for the weekend.”
L1 is "response-level." A smartphone can understand a simple, clear command, invoke a specific tool, and complete a single task—such as opening the calendar or setting a reminder for Saturday morning. It’s like a finger triggered by voice: you tell it where to tap, and it taps there.
L2 is "tool-level." The phone can understand simple intentions, perform limited reasoning, and invoke predefined tools to complete one-step or multi-step tasks with clear paths. It can check the weather in Hangzhou, generate a travel itinerary, and keep the results within the current conversation; however, the workflow is largely pre-defined by the product, and memory primarily relies on short-term context—essentially equivalent to installing a DouBao app on your phone.
The L3 commonly mentioned by smartphone manufacturers is the current true target: "assistive" level—where the phone must understand more complete objectives, proactively ask follow-up questions when information is insufficient, break tasks into steps, and select and sequence tools based on context. It must also possess short-term and long-term memory, so preferences such as "don't take the early train" or "prefer hotels close to West Lake" continue to apply in future tasks.
The same phrase, "arrange a trip to Hangzhou," might, at L3, be broken down into confirming dates and budget, checking transportation and weather, filtering hotels, generating an itinerary, and adding it to the calendar. When payments, sensitive data, or irreversible actions are involved, decision-making authority is returned to the user.
The difference between L1 and L3 is not that AI's responses become more human-like, but that the phone takes on increasingly complete tasks.
Testing on mobile devices is more concrete than a demo poster. The device must be tested in at least three typical scenarios, with a success rate of at least 80% for tool invocation tasks at the target difficulty level. Each individual task is expected to be completed within 5 minutes, with a maximum of three attempts. On-device capabilities are tested in an offline environment, while cloud-device collaborative capabilities are tested in an online environment.
In addition, L3’s long-term memory must cover at least three types of information, such as user basic information, tool preferences, conversation history, frequently visited locations, scheduled tasks, or usage patterns. However, complex reasoning and planning can be handled in the cloud, and there is no requirement for L3 to possess continuous evolutionary learning capabilities.
This explains the relationship between models and levels. Large models are responsible for understanding language, recognizing images, and planning tasks—they are the source of the phone’s core capabilities. However, no matter how powerful a model is, if it’s confined to answering questions within a chat box, it still cannot make the entire phone L3. Conversely, a model with modest parameters could still accomplish L3 tasks if it effectively integrates with the operating system, tool interfaces, user memory, and cloud capabilities.
In addition, the list of mobile device model registrations announced on July 15 is not the same as L3 testing. Registration determines whether a generative AI service can be provided in compliance with regulations; L3 testing evaluates whether the model, once integrated into a specific smartphone, can work together with the system and applications to complete tasks. The former sets the compliance baseline for services, while the latter measures the terminal’s capability.
What truly takes the exam is not a single model, but an entire team composed of the model, system, applications, and permission mechanisms.
Additionally, you may have noticed that smartphone manufacturers commonly refer to L3 as the "highest level currently available," which differs from automotive autonomous driving standards.
But in fact, this is also a play on words: L3 is indeed the highest level for which clear requirements have been provided and testing can currently be conducted; however, it is not the highest level in this certification system. In the detailed standards, L4 is named “Collaborative Level,” but its specific requirements are directly marked as “TBD.” The GB/Z 177.3—2026 standard for smartphones and tablets currently only defines testing methods for L1, L2, and L3.
In other words, an L3-level AI agent is still an assistant: it can break down tasks, invoke tools, and remember preferences, but primarily operates around a single device and one user goal. When it comes to "collaboration," the scope naturally expands to include multiple agents, multiple devices, and multiple service entities.
Can your phone delegate booking to a travel agent, assign budgeting to a payment service, and have your car system, headphones, and calendar jointly adjust the plan? Who is responsible if a task is executed incorrectly? Can permissions granted to one agent be passed on to another? These questions lack definitive answers, making it difficult to create reproducible test cases for L4.
L4 is left empty, not due to insufficient model capability, but because the collaborative rules, responsibility boundaries, and security mechanisms are not yet mature.
The L1–L4 levels on mobile devices cannot be applied using the liability logic of autonomous driving. Autonomous driving levels focus on vehicle takeover of dynamic driving tasks and accident liability allocation; AI terminal levels measure task complexity and system capability. The numbers may look the same, but the exams are entirely different.
Therefore, any claim that packages an L3 phone as “Level 4 autonomous AI” can essentially be understood as marketing rhetoric. As for a true L4 phone, we must first wait for standards to define what “collaboration” means—and who hits the brakes when collaboration fails.
Do you still need to change your phone?
For average consumers, the most common question I hear about this standard is: As this grading system is gradually implemented, will we need to upgrade to a new phone to access better AI Agent capabilities?
If asked strictly according to national standards, when will the L3 phone be released? The answer is: It has already been released.
The initial list does not consist solely of unreleased new phones: the Huawei Pura X Max was already released in April, and three Lenovo Moto phones were unveiled in May. The list also includes products still within their release window, such as the Honor Robot Phone. Additionally, the more widely known second-generation DouBao Phone is not included in the relevant list.
But what consumers truly care about is not L3 in the certification sense—it’s something else: When will they be able to buy a phone that reliably understands goals, performs tasks across apps, and doesn’t fail at the last step?
Honor has provided a clearer timeline: the Robot Phone will launch with the AgenticOS kernel in August 2026, and the Magic9 series is planned to offer a preview version in the fourth quarter.
On the native Android side, we are also advancing the AppFunctions and UI automation framework to enable apps to expose functionalities to agents, with plans to share more details by late 2026.
Based on this assessment, the second half of 2026 will be the first window for the concentrated emergence of "consumer-grade L3 experiences"; however, if you expect it to become a stable, mainstream capability on smartphones like cameras or payments today, a more realistic milestone is 2027.
According to Counterpoint’s projections, generative AI smartphones will account for 52% of global shipments by 2027, and agent capabilities will continue to expand to flagship devices.
However, this still cannot be directly translated into the market share of GB L3, as the two definitions are not the same. But it does provide a timeline we can expect for smartphone generational shifts: watch for flagship models to pilot in 2026, and see whether the app ecosystem and mid-range devices can catch up by 2027.
Therefore, if your current phone’s performance and battery life are not noticeably problematic, there’s no need to upgrade solely for an L3 certificate.
Standards allow complex reasoning to be offloaded to the cloud, and many capabilities may also be added to existing flagship devices through system updates; what truly determines the experience is whether manufacturers provide ongoing system maintenance and whether popular apps are willing to open their tool interfaces.
If you were already due for a device upgrade, the L3 can serve as a new criterion—but you should still consider three key factors: actual task success rate, coverage of commonly used applications, and whether each step is visible, pauseable, and reversible.
The certificate proves your phone has crossed a line, but it doesn't prove it's worthy of taking over your digital life.
What’s truly worth waiting for is the day when we no longer care about its rank, but simply remember that after saying, “Please take care of this,” the task was reliably completed.
