Tencent released and open-sourced Hy4 preview — a 770B-parameter, 49B-active Mixture-of-Experts model (78 layers: one dense layer plus 77 MoE layers with 256 routed experts and 1 shared expert, top-8 routed experts activated per token) with a 1M-token context window, released under Apache 2.0. It's available on Hugging Face, GitHub, ModelScope, AtomGit, Tencent Cloud, and OpenRouter, and free to use for two weeks at launch through Tencent's WorkBuddy and CodeBuddy products.
The generation-over-generation jump is real and large
Tencent's own comparison table (which we read directly rather than the highlighted bar-chart subset) shows the gap between Hy4 preview and its own predecessor, Hy3 preview, is genuinely substantial across nearly every listed benchmark: DeepSWE 28.0 → 64.3, ProgramBench 3.0 → 17.5, MathArena Apex 2025 38.7 → 74.2, BrokenArXiv 26.7 → 54.6, SWE-Marathon 5.0 → 31.9. Tencent calls this "the largest generation-over-generation gain we've measured," and on the numbers published, that specific claim holds up directly.
Counting the full table changes the "open-source frontier" framing
Tencent's announcement states this puts Hy4 preview "at the open-source frontier," and highlights 12 benchmarks in bar-chart form comparing it against Hy3, Qwen 3.8 Max, DeepSeek V4 Pro, GPT 5.6 Sol, GLM 5.3, Kimi K3, and Claude Opus 5. Worth counting directly rather than reading the chart's framing at face value: across the roughly 43 comparable rows in Tencent's own full table (spanning agentic coding, agentic search, working-agent, STEM, and reasoning categories), Hy4 preview posts the single highest score in only one row — SWE Atlas Codebase Q&A, at 64.0. In the 12 specific benchmarks Tencent chose to chart, Hy4 preview does not hold the top score in any of them per the same table's own numbers: GPT-5.6 Sol or Claude Opus 5 leads most of them (ProgramBench, SWE Atlas Refactoring, Agents' Last Exam, APEX-Agents, OneMillionBench, Humanity's Last Exam, HorizonMath), and other open-weight models take several more (GLM 5.3 and Kimi K3 both edge ahead on Terminal-Bench 2.1 and DeepSWE). That doesn't mean Hy4 preview is uncompetitive — it's frequently close behind the leader, and clearly ahead of Hy3 across the board — but "open-source frontier" reads more accurately as "near the frontier of what open models can do" than as "leads the field," once the table is counted row by row rather than read from the selected charts.
A methodology detail worth crediting
Tencent's own footnotes disclose that scores marked with an asterisk in its comparison table are from Tencent's own testing of the other models, not necessarily those vendors' own reported numbers — a specific, checkable distinction between self-reported and independently-reproduced figures that's easy to blur in a comparison table and worth naming directly since Tencent discloses it rather than leaving it ambiguous.
The blind human evaluation shows a narrow, not dominant, edge
Tencent ran a blind side-by-side evaluation: 163 internal experts rated model outputs on 203 real engineering tasks. Hy4 preview scored 2.99 average against GLM 5.3's 2.92 (46.8% wins, 12.8% ties, 40.4% losses) and Kimi K3's 2.94 (51.2% wins, 7.9% ties, 40.9% losses). Worth reading those win/loss splits directly: a 46.8%-to-40.4% margin is a real edge, not a rounding error, but it's also a near-even split with meaningful losses in both directions — a more honest picture of "narrowly ahead" than "wins."
A self-reported research case study, not independently verified
Tencent describes Hy4 preview managing several Codex sessions in parallel during a post-training research task, acting as the coordinating "researcher" directing Codex's exploration, and beating Codex working alone across all 8 evaluation targets in that exercise. This is Tencent's own internal case study relayed in the announcement, not an independently reproduced result — worth treating as illustrative of an intended capability rather than an externally verified claim.
What Tencent discloses as known issues
The announcement states directly that Hy4 preview currently spends "longer than necessary reasoning through complex tasks" and has "a tendency to over-verify its own work," and frames the release explicitly as an early version shipped to gather feedback, following the same approach used for Hy3 preview.
Pricing
Tencent's API pricing per 1M tokens: $0.042 cached input, $0.834 input, $2.501 output.
References: Tencent — Introducing Hy4 preview, read directly in full, including the full benchmark table · Hy4 preview on Hugging Face · Tencent Hunyuan's announcement on X · architecture and licensing details relayed from WebSearch summaries of coverage by Pandaily and BigGo Finance — this environment could not independently fetch those articles (egress restrictions) · related coverage: Frontier Arcade: trends & predictions