SemiAnalysis found public model-specific safety results for just 31 of 857 Chinese AI releases, with only nine available at or before launch.

Chinese AI Firms Published Safety Results for Just 3.6% of Model Releases

The420 Web Correspondent
6 Min Read

China’s leading artificial intelligence developers publicly released model-specific safety-test results for just 3.6% of their AI model launches over the past five years, according to a new study that raises questions about transparency as increasingly powerful systems reach the market.

California-based research firm SemiAnalysis examined 857 model releases between 2021 and September 15, 2026 from nine major Chinese AI companies.

The companies reviewed included Alibaba, ByteDance, Tencent, Baidu, DeepSeek, Moonshot AI, Z.AI, MiniMax and StepFun.

Researchers found published safety results for only 31 releases. Just nine models had those results available at or before launch.

FCRF Launches CP-FRM to Build India’s Next Generation of Fraud Risk Professionals

Most Models Had No Public Safety Disclosure

SemiAnalysis found no model-specific public safety disclosure for 813 of the 857 releases it examined.

That does not mean those models were never tested.

Companies may have conducted internal evaluations without publishing the results.

The study therefore measures transparency rather than proving the absence of safety work.

SemiAnalysis counted a disclosure only when specific safety results could be tied to an identifiable model.

General statements saying a model had been safety-trained or evaluated were not enough.

The tests considered included areas such as harmful outputs, resistance to jailbreaks, privacy, toxicity, refusal behaviour and dangerous capabilities.

Only Nine Models Had Results Available at Launch

The timing of safety disclosure was even more limited.

Only nine of the 857 models reviewed had public safety results available at or before the model was released.

That represents roughly 1.1% of all releases.

Some companies published evaluations later, but SemiAnalysis found that safety reporting frequently lagged behind deployment.

That matters because safety tests are most useful to users, researchers and regulators before or at the point when a powerful model becomes widely available.

Publishing results months later reduces their value for launch-time risk assessment.

Big Tech Companies and Startups Both Examined

The report covered both China’s largest internet companies and some of its fastest-growing AI startups.

Alibaba develops the Qwen model family.

ByteDance, Tencent and Baidu are also investing heavily in large language models and AI agents.

DeepSeek has attracted global attention with powerful open models, while Moonshot, Z.AI, MiniMax and StepFun have become major competitors in China’s rapidly expanding generative AI market.

The scale of the review is significant because these developers collectively account for a large share of China’s most advanced commercial and open AI releases.

China’s Rules Focus More on Applications

The findings come as China continues to build a national framework for AI governance.

Current rules place significant emphasis on content controls, security reviews, data protection and risks created by deployed AI applications.

But Reuters reported that China does not currently impose a broad mandatory system requiring frontier AI developers to publicly disclose model-level dangerous-capability evaluations before release.

That differs from some leading US AI companies, which increasingly publish system cards or safety reports describing how advanced models performed in areas such as cybersecurity, biological risk, deception and autonomous behaviour.

Those disclosures are also voluntary in many cases and vary significantly in detail.

Autonomous AI Makes Safety Testing More Important

The debate is becoming more urgent as AI systems move beyond generating text and begin operating tools, writing code and completing multi-step tasks.

An ordinary chatbot producing a harmful answer creates one type of risk.

An autonomous agent with access to software systems, credentials or network tools can potentially cause much greater damage.

Recent AI safety research has increasingly focused on whether advanced systems can conduct cyber operations, evade restrictions or behave deceptively during testing.

That makes independent scrutiny of safety evaluations more important as models become more capable.

Transparency Does Not Automatically Mean Safety

Publishing an evaluation does not prove that a model is safe.

Companies choose which tests to run, how results are measured and what information becomes public.

Safety benchmarks can also become outdated as models improve.

But disclosure allows outside researchers, businesses and governments to better understand what developers tested before deployment and where weaknesses remain.

Without model-specific results, users largely have to rely on company assurances.

That transparency gap is the central issue highlighted by the SemiAnalysis study.

Global Pressure for AI Safety Disclosure Is Growing

The report arrives as governments around the world debate how much evidence AI companies should be required to provide before releasing advanced models.

The European Union has created extensive obligations for general-purpose AI under its AI Act, including risk-management requirements for the most powerful systems.

Other governments are considering standards around testing, incident reporting and dangerous-capability evaluations.

China has also published broader AI safety-governance frameworks, but the SemiAnalysis findings suggest that public model-level reporting remains limited.

As competition accelerates, regulators face a difficult balance between encouraging rapid innovation and ensuring that increasingly capable AI systems are tested before widespread deployment.

What this means for you

The 3.6% figure does not prove that Chinese AI developers are releasing models without testing them. It shows that users and researchers rarely get to see model-specific safety evidence, making it much harder to independently judge the risks of increasingly powerful AI systems.

Follow for daily updates on cybercrime, corporate fraud, DFIR, hacking, investigations, and digital forensics

Stay Connected