China’s leading artificial intelligence developers publicly released model-specific safety-test results for just 3.6% of their AI model launches over the past five years, according to a new study that raises questions about transparency as increasingly powerful systems reach the market.
California-based research firm SemiAnalysis examined 857 model releases between 2021 and September 15, 2026 from nine major Chinese AI companies.
The companies reviewed included Alibaba, ByteDance, Tencent, Baidu, DeepSeek, Moonshot AI, Z.AI, MiniMax and StepFun.
Researchers found published safety results for only 31 releases. Just nine models had those results available at or before launch.
FCRF Launches CP-FRM to Build India’s Next Generation of Fraud Risk Professionals
Most Models Had No Public Safety Disclosure
SemiAnalysis found no model-specific public safety disclosure for 813 of the 857 releases it examined.
That does not mean those models were never tested.
Companies may have conducted internal evaluations without publishing the results.
The study therefore measures transparency rather than proving the absence of safety work.
SemiAnalysis counted a disclosure only when specific safety results could be tied to an identifiable model.
General statements saying a model had been safety-trained or evaluated were not enough.
The tests considered included areas such as harmful outputs, resistance to jailbreaks, privacy, toxicity, refusal behaviour and dangerous capabilities.
Only Nine Models Had Results Available at Launch
The timing of safety disclosure was even more limited.
Only nine of the 857 models reviewed had public safety results available at or before the model was released.
That represents roughly 1.1% of all releases.
Some companies published evaluations later, but SemiAnalysis found that safety reporting frequently lagged behind deployment.
That matters because safety tests are most useful to users, researchers and regulators before or at the point when a powerful model becomes widely available.
Publishing results months later reduces their value for launch-time risk assessment.
Big Tech Companies and Startups Both Examined
The report covered both China’s largest internet companies and some of its fastest-growing AI startups.
Alibaba develops the Qwen model family.
ByteDance, Tencent and Baidu are also investing heavily in large language models and AI agents.
DeepSeek has attracted global attention with powerful open models, while Moonshot, Z.AI, MiniMax and StepFun have become major competitors in China’s rapidly expanding generative AI market.
The scale of the review is significant because these developers collectively account for a large share of China’s most advanced commercial and open AI releases.
China’s Rules Focus More on Applications
The findings come as China continues to build a national framework for AI governance.
Current rules place significant emphasis on content controls, security reviews, data protection and risks created by deployed AI applications.
But Reuters reported that China does not currently impose a broad mandatory system requiring frontier AI developers to publicly disclose model-level dangerous-capability evaluations before release.
That differs from some leading US AI companies, which increasingly publish system cards or safety reports describing how advanced models performed in areas such as cybersecurity, biological risk, deception and autonomous behaviour.
Those disclosures are also voluntary in many cases and vary significantly in detail.
Autonomous AI Makes Safety Testing More Important
The debate is becoming more urgent as AI systems move beyond generating text and begin operating tools, writing code and completing multi-step tasks.
An ordinary chatbot producing a harmful answer creates one type of risk.
An autonomous agent with access to software systems, credentials or network tools can potentially cause much greater damage.
Recent AI safety research has increasingly focused on whether advanced systems can conduct cyber operations, evade restrictions or behave deceptively during testing.
That makes independent scrutiny of safety evaluations more important as models become more capable.
Transparency Does Not Automatically Mean Safety
Publishing an evaluation does not prove that a model is safe.
Companies choose which tests to run, how results are measured and what information becomes public.
Safety benchmarks can also become outdated as models improve.
But disclosure allows outside researchers, businesses and governments to better understand what developers tested before deployment and where weaknesses remain.
Without model-specific results, users largely have to rely on company assurances.
That transparency gap is the central issue highlighted by the SemiAnalysis study.
Global Pressure for AI Safety Disclosure Is Growing
The report arrives as governments around the world debate how much evidence AI companies should be required to provide before releasing advanced models.
The European Union has created extensive obligations for general-purpose AI under its AI Act, including risk-management requirements for the most powerful systems.
Other governments are considering standards around testing, incident reporting and dangerous-capability evaluations.
China has also published broader AI safety-governance frameworks, but the SemiAnalysis findings suggest that public model-level reporting remains limited.
As competition accelerates, regulators face a difficult balance between encouraging rapid innovation and ensuring that increasingly capable AI systems are tested before widespread deployment.
What this means for you
The 3.6% figure does not prove that Chinese AI developers are releasing models without testing them. It shows that users and researchers rarely get to see model-specific safety evidence, making it much harder to independently judge the risks of increasingly powerful AI systems.
Follow for daily updates on cybercrime, corporate fraud, DFIR, hacking, investigations, and digital forensics