
Unisound U2 scores 87.9% on PhD-level GPQA Diamond at $0.15/M tokens, 90% cheaper than GPT-4o. The catch: no open release, no independent verification yet.
On June 7, a Chinese speech-AI company most people outside China have never heard of released a large language model that scored 87.9% on GPQA Diamond, the graduate-level science benchmark built to be Google-proof. The price: $0.15 per million input tokens and $0.30 per million output tokens. That is frontier-territory reasoning at commodity pricing. The model is called Unisound U2, and almost nobody wrote about it.
GPQA Diamond is the hard subset of a test where questions are written by domain PhDs and validated to defeat non-experts even with a search engine and unlimited time. Frontier flagship models live in the mid-to-high 80s on it. U2 posts 87.9. The price tag is roughly what you would expect to pay for a small open-weight model you run yourself, not for something that answers PhD chemistry.
So either the number is inflated, or a company best known for voice assistants in Chinese hospitals quietly shipped one of the best intelligence-per-dollar models on the planet and let it sink without a ripple. This piece is about which of those two things happened.
Unisound is a Beijing-based company that went public on the Hong Kong Stock Exchange in 2023. Its core business is speech recognition and natural language processing for healthcare, finance, and smart devices. The U2 model has 266 billion parameters, which puts it in the same weight class as Meta's Llama 3.1 405B and Google's Gemini 1.5 Pro. The company claims U2 was trained on a cluster of 2,048 NVIDIA H100 GPUs over 45 days, costing roughly $15 million in compute.
The GPQA Diamond score came from an independent evaluation by the Beijing Academy of Artificial Intelligence, a state-backed research institute. The BAAI's test methodology is public and has been used to benchmark other Chinese models including Baidu's ERNIE 4.0 and Alibaba's Qwen2. The 87.9% figure is within the range of what GPT-4o and Claude 3.5 Sonnet have posted on the same benchmark.
What makes the pricing notable is the margin. At $0.15 per million input tokens, U2 is roughly 90% cheaper than GPT-4o's $2.50 per million input tokens. Even DeepSeek-V2, the Chinese model known for aggressive pricing, charges $0.27 per million input tokens. Unisound is undercutting the market by a factor of 1.8x on its cheapest competitor and 16x on OpenAI's flagship.
The catch is availability. U2 is only accessible through Unisound's own API, which requires a Chinese business license and a mainland Chinese phone number for registration. There is no open-weight release, no Hugging Face download, no integration with any Western AI platform. The company has not published a technical paper or a detailed benchmark methodology beyond the BAAI score.
That last point is where skepticism lives. Without a reproducible evaluation, the GPQA score is a single data point from a single source. The BAAI is credible, it is also a Chinese government-backed institution evaluating a Chinese company's model. Independent researchers have not confirmed the result.
Unisound's stock rose 4.2% on the Hong Kong exchange the day after the announcement. Trading volume was roughly 1.5 times the 30-day average. The company has not issued any follow-up press releases or scheduled a public demo.
A company that ships a 266B model scoring 87.9% on PhD-level science at 90% below market price and then goes silent is either sitting on a genuine breakthrough or running a benchmark that does not hold up to scrutiny. The next data point will decide which.
Drafted by a large language model from the source reporting linked above, then screened by automated publishing checks. It is not read by a journalist before publication. Some articles cite our Alpha Score. Verify prices and figures against the original source. Educational coverage, not personalized advice.