Controlling Reasoning Effort in LLMs
- It has been almost two years since OpenAI released o1, a model that popularized the idea of LLM-based reasoning models.
- DeepSeek-R1 followed about four months later, together with details of a reinforcement learning with verifiable rewards (RLVR) recipe to train such reasoning models.
- Last week, OpenAI released the GPT-5.6 model family.
Unverified
- It has been almost two years since OpenAI released o1, a model that popularized the idea of LLM-based reasoning models.
- DeepSeek-R1 followed about four months later, together with details of a reinforcement learning with verifiable rewards (RLVR) recipe to train such reasoning models.
- Last week, OpenAI released the GPT-5.6 model family.
Sources: Sebastianraschka