Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It is leaps and bounds better than LLMs. For one you are doing RL which is classic AI like tuning that optimizes a reward function with nice qualities- it's the same stuff used to train chess games and Go by showing them the actual moves.

LLMs pre o1 and deepseek R1 were RHLF tuned which is like if you trained a LM how to play chess by showing people two boards and doing a vibe check on which "looks" better.

Think of it this way say you were dropped in a maze that you had to solve but you could do only one of two things:

1. Look at two random moves from your start position and selected which one looked better to get out.

2. Made a series of moves and then backtracked, then use a quantitative function to exploit the best path.

The latter is what R1 does and it chooses optimal and more certain path to success.

Apply this to math and coding tokens and you have a competitive LLM.



I am using the 32Gb distilled model on my local 3090 with Continue in VSCode. It beats everything out of the water.


How many tokens/s do you get on a 3090? With the extra tokens for the internal monologue, is it still performant enough for smooth VSCode integration?


Any idea how to use a cloud hosted version with cursor?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: