Diffusion-Proof is the first framework to train and apply diffusion language models for formal theorem proving. It introduces dLLM-Prover-7B for whole-proof writing with long-range coherence and dLLM-Corrector-7- for local proof correction using bidirectional information. The framework outperforms auto-regressive LLM baselines by 1.61% on ProofNet-Test and 6.14% on MiniF2F-Test, and solves an IMO problem beyond the capability of DeepSeek-Prover-V2-7B.
Diffusion-Proof: First Framework for Diffusion LLMs in Formal Theorem Proving
Lean as Process-Verified Reward Oracle in RL for Theorem Proving
This work shows that Lean can serve as a symbolic process oracle, providing fine-grained, verified feedback during reinforcement learning. By parsing proof attempts into tactic sequences and using Lean's elaboration to mark sound steps and first failures, the system generates dense, type-theoretic reward signals. Experiments demonstrate tactic-level supervision outperforms outcome-only methods on benchmarks like MiniF2F and ProofNet, highlighting Lean's role as both evaluator and training reward source.
Frustrated Synchronization Network Outperforms Transformers
The Frustrated Synchronization Network (FSN) achieves lower validation loss than a RoPE-SwiGLU transformer at every epoch on character-level text and code tasks. At one million parameters, FSN converges to a validation loss of 1.5953 ± 0.0014, outperforming the transformer's converged loss of 1.611. This advantage persists up to four million parameters, with ongoing evaluations beyond that scale.
Frontier Post-Training Recipe Review with Finbarr Timbers
The podcast reviews the evolution of post-training recipes in large language models, from InstructGPT to 2026 frontier models. It highlights Multi-Teacher On-Policy Distillation (MOPD) as the dominant pattern, where domain-specialist models are trained and then distilled into a general student model via on-policy distillation, scaling to over 10 teachers in models like DeepSeek V4 and Nemotron 3 Ultra.
Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
A comparative analysis evaluates the capabilities of Claude 3.7 Sonnet, Claude 3.5 Sonnet, OpenAI o3-mini, DeepSeek R1, and Grok 3 Beta across math, coding, and reasoning benchmarks.
Anthropic releases Claude 3.7 Sonnet with dynamic reasoning control
Anthropic has released Claude 3.7 Sonnet, a model designed for real-world business and developer use cases rather than just benchmark optimization. The release introduces the ability to dynamically switch between standard and advanced reasoning modes, allowing users to control the number of tokens allocated to "thinking time."