LLMs Compared - Which is better

Btw gemini 3.1 also became worse at complex problems I think. It got better though at doing something similar to what I’m saying as oppose to the complete opposite :joy:

1 Like

Say that clearly please?

Is it better or worse at coding?

I might be wrong but in my first few tries with it I find that it’s not as bad as it used to be in terms of ignoring clear prompt and doing the exact opposite.

I also think it got worse at actual training through complex issues and now it’s really looking for shortcuts

1 Like

Unsurprisingly Copilot ranks worst in public opinion.

1 Like

I don’t use LLMs too much, but I jump between Gemini ChatGPT and Grok, sometimes I get what I want from one of them, usually the experience is very frustrating, basically talking to a know-it-all who knows nothing, and is always confidently saying false things.

I recently had a very simple task I gave all 3, I did a video project for the same client for a bunch of years, most years took 40-45 hours, but this year took 65 hours, I wanted them to analyze my pdfs I exported from my time tracker stating what I worked on when, and compare all the years to see if there are any mistakes, if not what the reason for the extra time is. They all started out looking pretty good, until I realized they were making up data, nothing like the actual data I gave them. sometimes math was wrong too. Even when I corrected them they just made up other things, and I obviously don’t want to go through the data myself, that’s why I’m using AI.

Then I decided to try Claude. Sonnet 4.6 extended. I caught it on hallucinated data at first, but the difference is what it did when I pointed out that I need to make sure there are no more mistakes. It took a while, but it extracted and parsed the data from the pdfs, using python scripts, and intelligently correcting issues in its data BEFORE giving it to me,

and actually was completely accurate wherever I checked from then on. (Turns out even where I thought the others were right, they were wrong)

So as of now I like what I see in Claude (except I was coming up on the limit), and I hope it continues…

5 Likes

https://marketplace.microsoft.com/en-us/product/saas/wa200009404?tab=overview

Testing it… So far so good, hasn’t sped me up a crazy amount yet, don’t fully trust it yet.

I also deal alot with vba and I’m not sure how it handles that, if it does.

I used it and am very happy, I even helped a user that does all of the finances for a company and had a sheet and the data was not matching the total amount from QuickBooks. Less than a minute it’s found more than 15 mistakes and cells containing raw data instead of formulas.

And this was after the user tries for a few hours to manually fix everything.

1 Like

Awesome will give it more of a go. I hope it does vba within excel I will check later today.

1 Like

just hit my limit on claude

im assuming your using a $20 plan or free…

no paid plan

I agree now! :rofl: I thought it would be just like the upgrade from gemini free to gemini pro which isn’t worth the money if you’re not a spender but Claude is totally different. The free sonnet is amazing and beats the others but opus in the terminal is insane!!

It’s only my first day but So far so good! (Which is why I’m still up now :rofl::yawning_face:)

5 Likes

Gemini web just got a major UI refresh.

1 Like

As well as lower limits.

Looks very AI generated

True, but unfortunately, AI-generated content is becoming today’s norm.

1 Like

Is there a problem with that?

Yes and that’s acceptable, but not for a trillion dollar company. I could use their OWN Google Stitch to produce something better

What is google stitch?

Sometimes