Chinese AI Safety Reports Covered Just 31 of 857 Model Releases

Chinese AI companies published safety-test results for just 3.6% of their AI model releases, according to a new report from research firm SemiAnalysis. The findings raise concerns about transparency as companies continue to develop and release new AI systems.

The researchers reviewed 857 AI model releases from nine leading Chinese developers between 2021 and September 15, 2026. Only 31 releases had publicly available safety results that could be linked to a specific model.

The report also found that just nine releases, or 1.1% of the total, had safety-test results available at or before launch. The findings suggest that public safety reporting has not kept pace with the number of models being released.

However, the results do not prove that companies failed to test their models. Some developers may have conducted safety checks without publishing the findings.

Nine Major Chinese AI Firms Were Included in the Study

SemiAnalysis examined AI model releases from nine major Chinese developers. These included Alibaba, ByteDance, Tencent, Baidu, DeepSeek, Moonshot, Z.AI, MiniMax and StepFun.

The study covered 857 releases during the review period. This included 741 product models and 116 research models. The researchers checked public documents, including model cards, release notes and technical reports. They looked for safety-test results linked to specific models.

General statements about safety training did not qualify. Companies needed to provide actual evaluation results that could be connected to a particular model. This distinction matters because a company may claim that it prioritises AI safety without publishing evidence of how a model performed during testing.

The report found that 813 releases had no public safety disclosures in the materials reviewed. This accounted for about 94.9% of all releases.

Most AI Models Launched Without Public Safety Test Results

Most AI Models Launched Without Public Safety Test Results

The timing of safety reports was another concern. Of the 857 releases, only nine had published safety evaluations available at or before launch. This means developers publicly shared relevant test results for very few models before making them available.

Some companies published results after releasing their models. According to the report, 16 releases had safety results published later. The median delay was 42 days. In some cases, the delay was much longer. The report recorded a maximum delay of 349 days for a published safety result.

This gap can make it harder for users and developers to assess a model’s risks before adopting it. It can also limit the information available to businesses that want to use these systems in their products.

However, a delayed report does not automatically mean a model was unsafe at launch. It shows that the public did not have access to the relevant findings at the same time.

What Do AI Safety Evaluations Actually Measure?

AI safety evaluations help researchers understand how a model behaves in different situations. These tests can examine whether a system produces harmful responses, reveals private information or follows instructions that it should reject.

Researchers may also test whether a model can resist jailbreak attempts. These attempts use carefully written prompts to bypass a system’s safety restrictions. Other evaluations examine whether advanced models can support dangerous activities or show risky behaviour when given complex tasks.

Publishing these results helps users understand a model’s limits. It also allows outside researchers to review the findings and identify possible weaknesses.

Without detailed results, users may have to rely on a developer’s general safety claims. This makes it harder to compare models based on publicly available evidence.

ALSO READ: OpenAI Reveals 6 Alarming AI Model Behaviors in New Safety Reports

China’s AI Safety Framework Faces Questions Over Public Testing

China's AI Safety Framework Faces Questions Over Public Testing

The findings have renewed questions about how China regulates advanced AI development. China has introduced rules and guidance covering AI services and the risks they may create. However, the report highlights concerns about the lack of mandatory public, model-specific safety evaluations for the releases examined.

This is important as AI systems become capable of completing more tasks with limited human involvement. Such systems, often called AI agents, can interact with tools and carry out multiple steps to complete a task.

If an agent behaves unexpectedly, the consequences may extend beyond an incorrect answer. Depending on its permissions, it could also take actions that affect users, businesses or computer systems.

Public safety reports can help developers and customers understand these risks before deploying a model. They can also support independent research into how well safety measures work.

The report does not establish that Chinese AI models are less safe than every competing model. Instead, it identifies a gap in publicly available evidence about their safety performance.

AI Industry Faces a Growing Gap Between Releases and Safety Reports

The AI industry is releasing new models at a rapid pace, but detailed public information about their safety performance remains limited. This gap makes it harder for businesses and users to assess potential risks before adopting new systems.

For businesses, this creates an additional challenge when choosing AI systems. A model’s performance, price and speed may be important, but its safety record also matters. Companies that use AI models in customer service, software development or other sensitive tasks need to understand their limitations. Published evaluations can help them make more informed decisions.

Developers can also improve transparency by sharing model-specific test results before launch. They can explain what they tested, what problems they found and what steps they took to address them.

The SemiAnalysis report does not show that every model without published results was released without testing. However, it shows that detailed public safety evidence was available for only a small share of the Chinese AI releases examined.

As AI systems become more capable, the gap between model releases and public safety reporting is likely to remain an important issue for developers, businesses and regulators.

This entry was posted in AI News. Bookmark the permalink.

Leave a Reply

Your email address will not be published. Required fields are marked *