Gemini 3.7 Flash takes the top spot this week after shipping just three weeks after its predecessor, while the introduction of the LMArena Expert framework reshuffled leaderboard positions across the board.

This Week at a Glance

Gemini 3.7 Flash from Google DeepMind claims the number one position following its release within three weeks of Gemini 3.6 Flash, generating high volume across developer-community discussion and media coverage. The single largest structural shift on the board comes from the LMArena Expert evaluation framework, which introduced a new leaderboard tier that altered Elo rankings for multiple frontier models.

Heat Index

Rank Model Developer Change Heat Index Why It's Here
1 Gemini 3.7 Flash Google DeepMind New 92 Released three weeks after 3.6 Flash; 3.5 Pro remains delayed (press report)
2 LMArena Expert Framework LMSys / Arena.ai New 88 New evaluation framework identifying expert-level prompts; powers Expert leaderboard (official blog)
3 Qwen3-235B Alibaba Qwen 82 Leads open-weight ecosystem per Hugging Face mid-2026 report (official report)
4 DeepSWE Not disclosed in sources New 76 Scored 62.7; supports thinking-intensity control and native Responses API (community testing)
5 Sentence-Embedding Model Not disclosed in sources New 65 Reached nearly 1.6 billion pulls on Hugging Face (official report)

Reading the Board

Gemini 3.7 Flash occupies the first rank because it shipped rapidly. According to press reports tracking the September 2026 release wave, Google DeepMind launched Gemini 3.7 Flash only three weeks after introducing Gemini 3.6 Flash. The same reports note that the flagship Gemini 3.5 Pro, originally scheduled for a June push, remains delayed. This rapid iteration cycle drove sustained media coverage and developer-community discussion throughout the seven-day window, securing its position without requiring a new benchmark publication this specific week. (press report)

The second entry is not a large language model (LLM) but an evaluation infrastructure update that moved the entire field. Arena.ai officially introduced Arena Expert, a new LMArena evaluation framework designed to identify the toughest, most expert-level prompts from real users. This framework powers a new Expert leaderboard. Because LMArena re-ranks roughly weekly as votes accumulate, this structural addition altered how models are compared, generating significant discussion on X, Reddit, and Hacker News among practitioners who rely on these Elo scores for model selection. (official blog)

Alibaba Qwen holds the third position based on ecosystem dominance rather than a version update this specific week. Hugging Face's mid-2026 ecosystem report documents a structural split in open-weight AI, noting that trillion-parameter Chinese models grab headlines, with Qwen leading the open-weight category. While the report covers the broader mid-year period, the ongoing discussion around Qwen3-235B maintained high engagement velocity on Hugging Face trending trackers through this week, keeping it on the board. (official report)

DeepSWE enters at number four based on community testing results circulated this week. Community testing shows DeepSWE achieved a score of 62.7, alongside support for thinking-intensity grading control and a native Responses API. Analysis from community testers focused on its engineering transition from a long-text reader to an independent delivery engineer, comparing its performance alignment and pricing against head-tier mid-range flagships. No official developer announcement or parameter count was published in the available material for this specific week. (community testing)

A small sentence-embedding model rounds out the board at number five. Hugging Face's 2026 Open Model Report highlights that while trillion-parameter models dominate headlines, this specific embedding model accumulated nearly 1.6 billion pulls. This download volume represents a distinct type of heat—infrastructure adoption rather than frontier benchmark competition—justifying its inclusion despite lacking a traditional LLM architecture or a release event this week. (official report)

Open Weights vs. Closed

This week's board contains three entries tied to open weights: Qwen3-235B, the sentence-embedding model, and DeepSWE, which community testers evaluated using open-weight methodologies. The highest rank held by an open-weights model is three, occupied by Qwen3-235B.

The closed camp is represented solely by Gemini 3.7 Flash at number one. The remaining slot belongs to the LMArena Expert framework, which applies to both camps. Hugging Face's mid-2026 ecosystem report explicitly documents a structural split in open-weight AI, observing that trillion-parameter Chinese models generate headlines while smaller utility models accumulate higher raw pull counts. This divergence between hype and actual deployment volume defines the current open-weights landscape.

Watch Next Week

  • LMArena Expert leaderboard vote accumulation: The new Expert leaderboard re-ranks roughly weekly as votes accumulate, meaning next week's snapshot will reflect the first full cycle of expert-prompt evaluations. (official blog)
  • As of publication, no major releases have been officially announced for next week.

Gemini 3.7 Flash leads the September 8–14 board due to its rapid release cadence, while the LMArena Expert framework represents the largest structural shift in evaluation this week. The Heat Index is an editorial estimate based on public information. It is not official data from any platform and is not a recommendation for model selection or investment.


参考资料

  1. LLM News Today (September 2026) - AI Model Releases
  2. LLM Updates (September 2026) - AI Model Releases ...
  3. AI Updates Today (September 2026) - Latest AI Model Releases
  4. New AI Models: Live Release Tracker (Last 24 Hours) | BenchLM
  5. LLM Releases — the model release tracker
  6. AI/TLDR — New AI Models, Tools & Papers This Week
  7. AI Model Release Tracker: New LLM Launches, Updated Weekly | AI Flash ...
  8. Hugging Face's 2026 Open Model Report: Qwen Leads, Hype vs. Reality
  9. The Open Weights — open-source AI, tracked daily
  10. Trending open-source AI models · The Open Weights