Is that a question or a statement?
It sounds like Gemini 3 is coming , and it will be big :
Google is likely to release Gemini 3 on October 9 after extensive testing on AI Studio. Early tests show significant performance on coding tasks.
1 Like
I didnât see anything in beta yet
I guess if Google didnât reach out to you to let you know about it then itâs probably not true.
AI
GEMINI 3 DROPS OCTOBER 9th!
Googleâs latest AI model is hitting AI Studio with major upgrades:
Massive coding performance boost - outperforms current Gemini 2.5 & Anthropic Sonnet 4.5
Better SVG generation capabilities
New âMy Stuffâ gallery for your AI creations
Tighter Google app integrations
Agent mode with browser control coming soon
Perfect timing for developers and tech creators ready to level up their AI workflows. Mark your calendars!
#Gemini3 #GoogleAI #AIStudio
Anything about it being released on October 9th is just internet hock with no real source or proof, just speculation.
Maybe yeah, maybe no, weâll wait and see.
Grok 5 is coming!
Elon Musk says Grok 5 will be released before the end of 2025 .
Iâve always been intrigued by this very common viewpoint. Personally, I cannot think of any personal data or info Iâd be uncomfortable with Xi Jinping accessing, but comfortable with Sam Altman accessing.
In other words, if itâs truly private, I will not send it to any hosted llm, and if itâs not, I donât care if China sees my stupid code I am trying to debug,
2 Likes
Btw Qwen and Qwen coder are both really good, and are open weight.
I hate to say it, but China is leading OpenAI and all US companies in the industry of open AI. (Meaning, open source AI. Sorry, I had to do it.
ars18
October 21, 2025, 5:01am
53
To self host? Thereâs definitely great free options but you have to have the hardware.
Iâve been using z.ai glm 4.6 as an add on to Claude code when I hit limits.
Cheapest plan is $3 a month. Pretty decent model
1 Like
If you want to go through the effort, itâs way cheaper to pay for the cloud compute to run a model than to pay for whatever these companies are chaging for these models.
ars18
October 21, 2025, 5:28am
55
I dunno⌠They are pretty cheap
Info and rankings of top self-hostable models
(I threw in ChatPT-5 and Claude 4.5 so you can see how the other models compare to it)
Here is a ranked, 1â10 scoring table across key coding metrics and aspects, including a ârelative to flagshipâ score that situates each model against GPTâ5 and Claude 4.5 on SWEâbench Verified and agentic coding tasks.[1][2]
Ranking table
Model
Open weights
SWEâbench Verified (%)
TerminalâBench
HumanEval pass@1
LiveCodeBench
Selfâhostability (1â10)
Agentic coding (1â10)
Relative to flagship (1â10)
GPTâ5 (high)
No [3]
72.5 [1]
n/a [1]
n/a [1]
n/a [1]
1 [3]
10 [1]
10 [1]
Claude Sonnet 4.5
No [2]
70.6 [1]
61.4 (OSWorld; agentic proxy) [2]
n/a [2]
n/a [2]
1 [2]
10 [2]
10 [1]
Qwen3âCoderâ480BâA35BâInstruct
Yes (Apacheâ2.0 per LB) [1]
67.0 [1]
n/a [1]
n/a [4]
n/a [4]
5 (very large MoE) [1]
9 (frontierâlevel reports) [5]
9 [1]
DeepSeekâV3.1
Yes [6]
66.0 [1]
n/a [1]
n/a [7]
Top performer claims (V3 family) [7]
6 (open, heavy MoE) [6]
9 (strong repoâlevel reports) [7]
9 [1]
Kimi K2
Yes [8]
65.4 [1]
25.0 [9]
n/a [10]
n/a [9]
6 (open, sizable) [8]
7 (agentic model family) [10]
9 [1]
GLMâ4.5
Yes [11]
64.2 [12]
37.5 [9]
n/a [12]
n/a [12]
6 (32B active MoE) [12]
9 (strong terminal automation) [9]
8 [1]
Llama 4 Maverick (17Bâ128E)
Yes (Llama license) [13]
21.04 [1]
n/a [1]
n/a [14]
n/a [14]
9 (17B active, efficient) [14]
4 (generalist, smaller) [1]
3 [1]
Falcon 2â11B
Yes [15]
n/a [15]
n/a [15]
29.6 [16]
n/a [15]
10 (lightweight dense) [16]
3 (older training mix) [15]
3 [16]
BLOOMâ176B
Yes [17]
n/a [17]
n/a [17]
Low vs modern code models [18]
n/a [19]
3 (very large dense) [19]
2 (not codeâtuned) [19]
1 [19]
How to read scores
The 1â10 ârelative to flagshipâ score is normalized primarily to SWEâbench Verified, where GPTâ5 and Claude 4.5 sit at the top of current public leaderboards, so models at 65â69% land near 9 and 70%+ at 10.[2][1]
âAgentic codingâ reflects repo/terminal automation indicators (e.g., TerminalâBench, OSWorld for agentic computer use), along with modelâfamily claims and evaluation writeâups.[9][2]
âSelfâhostabilityâ weighs open weights availability and effective active parameters; small dense or smallâactive MoE models score higher for local deployment feasibility.[12][13]
Key takeaways
Among openâweights, GLMâ4.5, DeepSeekâV3.1, Kimi K2, and Qwen3âCoder reach midâtoâhigh 60s on SWEâbench Verified, placing them within one tier of GPTâ5/Claude 4.5 for repositoryâlevel coding.[12][1]
GLMâ4.5 shows particularly strong terminal automation scores, and Kimi K2 appears in multiple comparisons as a competitive agentic coder for practical use.[9][1]
Llama 4 Maverick is an efficient 17Bâactive MoE thatâs easy to selfâhost, but its repoâlevel coding score is far below the top tier; Falcon 2 and BLOOM are best seen as lightweight or legacy baselines rather than primary coders.[16][1]
Sources
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
ars18
November 3, 2025, 3:07pm
58
Itâs a bit dumber then CC (plan plan plan) but context and limits are much bigger. Never hit a limit yet and I abused it.
Worth a try if you are dipping your toes in the water and donât want to pay much.
Perplexity Pro for $12
1 year subscription for Perplexity Pro for only $12.
I highly recommend it , I have Perplexity Pro, love it, and use it all the time.
I canât guarantee that this site is legit, youâll have to figure that out yourself, but it looks good on Trustpilot.
https://www.trustpilot.com/review/cheapgpt.store?languages=all
2 Likes
Flippy
November 11, 2025, 12:05am
60
You can get a free year of Perplexity Pro with a Venmo account.
You like it better than ChatGPT?