<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Finding Theta</title><description>finding-theta technical blog: notes on search, inference, optimization, and reproducible engineering.</description><link>https://www.findingtheta.com/</link><language>en</language><atom:link href="https://www.findingtheta.com/rss.xml" rel="self" type="application/rss+xml"/><item><title>Reinforce Tactics: A Technical Compendium and Analysis of Large Language Model Performance in Stochastic Strategy Environments</title><link>https://www.findingtheta.com/blog/reinforce-tactics-a-technical-compendium-and-analysis-of-large-language-model-performance-in-stochastic-strategy-environments/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/reinforce-tactics-a-technical-compendium-and-analysis-of-large-language-model-performance-in-stochastic-strategy-environments/</guid><description>Explore a technical case study on &apos;Reinforce Tactics,&apos; a new RL environment where traditional game theory outperforms the latest LLMs by a massive margin, exposing critical flaws in token-based reasoning.</description><pubDate>Thu, 01 Jan 2026 00:00:00 GMT</pubDate><category>projects</category><category>reinforcement-learning</category><category>reinforce-tactics</category><category>ppo</category><category>large-language-models</category></item><item><title>Advanced Architectures and Methodologies in Visual Reinforcement Learning: A Technical Analysis of the ViZDoom Platform</title><link>https://www.findingtheta.com/blog/advanced-architectures-and-methodologies-in-visual-reinforcement-learning-a-technical-analysis-of-the-vizdoom-platform/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/advanced-architectures-and-methodologies-in-visual-reinforcement-learning-a-technical-analysis-of-the-vizdoom-platform/</guid><description>Explores how the ViZDoom platform drives advancements in embodied AI, detailing the evolution from standard Deep Q-Networks to hierarchical architectures capable of mastering complex, 3D environments</description><pubDate>Mon, 01 Dec 2025 00:00:00 GMT</pubDate><category>experiments</category><category>reinforcement-learning</category></item><item><title>The Evolution of Imagination: A Deep Dive into DreamerV3 and its Conquest of Minecraft</title><link>https://www.findingtheta.com/blog/the-evolution-of-imagination-a-deep-dive-into-dreamerv3-and-its-conquest-of-minecraft/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/the-evolution-of-imagination-a-deep-dive-into-dreamerv3-and-its-conquest-of-minecraft/</guid><description>Explore DreamerV3, the AI that taught itself to find diamonds in Minecraft, revolutionizing reinforcement learning with its powerful world model.</description><pubDate>Sat, 01 Nov 2025 00:00:00 GMT</pubDate><category>deep-dives</category><category>reinforcement-learning</category><category>world-models</category></item><item><title>The Unseen Hand: Guiding a Virtual Drone with Sparse and Dense Rewards</title><link>https://www.findingtheta.com/blog/the-unseen-hand-guiding-a-virtual-drone-with-sparse-and-dense-rewards/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/the-unseen-hand-guiding-a-virtual-drone-with-sparse-and-dense-rewards/</guid><description>Exploring training a drone with reinforcement learning, focusing on the trade-offs between sparse and dense rewards to achieve agile flight while avoiding unintended &quot;reward hacking.&quot;</description><pubDate>Mon, 01 Sep 2025 00:00:00 GMT</pubDate><category>experiments</category><category>machine-learning</category><category>reinforcement-learning</category><category>robotics</category></item><item><title>Ultimate Guide to Contextual Bandits: From Theory to Python Implementation</title><link>https://www.findingtheta.com/blog/ultimate-guide-to-contextual-bandits-from-theory-to-python-implementation/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/ultimate-guide-to-contextual-bandits-from-theory-to-python-implementation/</guid><description>Discover the ultimate guide to contextual bandits, covering everything from core theory and key algorithms to a complete Python implementation with code for building powerful personalization and recommendation systems</description><pubDate>Fri, 01 Aug 2025 00:00:00 GMT</pubDate><category>deep-dives</category><category>contextual-bandits</category></item><item><title>Mastering Autonomy: A Comprehensive Guide to the AutoDRIVE Ecosystem and Reinforcement Learning</title><link>https://www.findingtheta.com/blog/mastering-autonomy-a-comprehensive-guide-to-the-autodrive-ecosystem-and-reinforcement-learning/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/mastering-autonomy-a-comprehensive-guide-to-the-autodrive-ecosystem-and-reinforcement-learning/</guid><description>A complete guide to the AutoDRIVE ecosystem. Learn to train self-driving car agents from scratch using reinforcement learning in this powerful Unity-based simulator.</description><pubDate>Tue, 01 Jul 2025 00:00:00 GMT</pubDate><category>experiments</category><category>robotics</category></item><item><title>Agentic AI: The Autonomous Evolution of Machine Learning and Its Dawn in Robotics</title><link>https://www.findingtheta.com/blog/agentic-ai-the-autonomous-evolution-of-machine-learning-and-its-dawn-in-robotics/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/agentic-ai-the-autonomous-evolution-of-machine-learning-and-its-dawn-in-robotics/</guid><description>Explore Agentic AI, the next frontier in machine learning. Discover how autonomous agents learn, act independently, and are revolutionizing the field of robotics.</description><pubDate>Sun, 01 Jun 2025 00:00:00 GMT</pubDate><category>deep-dives</category><category>robotics</category><category>agents</category></item><item><title>Serving Up Some Robotics: Setting Up a Tennis Environment in MuJoCo</title><link>https://www.findingtheta.com/blog/serving-up-some-robotics-setting-up-a-tennis-environment-in-mujoco/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/serving-up-some-robotics-setting-up-a-tennis-environment-in-mujoco/</guid><description>Build a MuJoCo robot tennis simulation! Learn to set up a wall tennis environment, tackle physics/control challenges, understand its architecture, and improve it for robotics or reinforcement learning projects with MuJoCo</description><pubDate>Thu, 01 May 2025 00:00:00 GMT</pubDate><category>projects</category><category>mujoco</category><category>courtside-dynamics</category><category>reinforcement-learning</category></item><item><title>Bridging Worlds: How Visual Language Action Models are Teaching Robots to See, Understand, and Act</title><link>https://www.findingtheta.com/blog/bridging-worlds-how-visual-language-action-models-are-teaching-robots-to-see-understand-and-act/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/bridging-worlds-how-visual-language-action-models-are-teaching-robots-to-see-understand-and-act/</guid><description>Understand how Visual Language Action Models (VLAs) let robots follow commands by fusing AI vision and language for action, demonstrated with a conceptual Python robotics simulation</description><pubDate>Tue, 01 Apr 2025 00:00:00 GMT</pubDate><category>deep-dives</category><category>robotics</category><category>vla</category></item><item><title>From Zero to Dino-Roar: Teaching a T-Rex to Walk with MuJoCo and Reinforcement Learning</title><link>https://www.findingtheta.com/blog/from-zero-to-dino-roar-teaching-a-t-rex-to-walk-with-mujoco-and-reinforcement-learning/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/from-zero-to-dino-roar-teaching-a-t-rex-to-walk-with-mujoco-and-reinforcement-learning/</guid><description>A deep technical walkthrough of building a biomechanically faithful T-Rex in MuJoCo and training it to balance, walk, and hunt through three-stage curriculum learning with PPO — achieving a 96.7% bite success rate on commodity GPU hardware. Covers MJCF model design, Gymnasium environment architecture, reward engineering with mathematical formulations, and real training results from the Mesozoic Labs open-source project.</description><pubDate>Sat, 01 Mar 2025 00:00:00 GMT</pubDate><category>projects</category><category>mujoco</category><category>mesozoic-labs</category><category>reinforcement-learning</category><category>ppo</category></item><item><title>Using Reinforcement Learning for Stock Trading with FinRL</title><link>https://www.findingtheta.com/blog/using-reinforcement-learning-for-stock-trading-with-finrl/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/using-reinforcement-learning-for-stock-trading-with-finrl/</guid><description>A walkthrough of reinforcement learning for stock trading with FinRL: the MDP and Bellman theory underneath it, building a trading environment, and training and backtesting agents with Stable-Baselines3.</description><pubDate>Sat, 01 Feb 2025 00:00:00 GMT</pubDate><category>experiments</category><category>reinforcement-learning</category></item><item><title>Mastering Robotic Manipulation with Reinforcement Learning: TQC and DDPG for Fetch Environments</title><link>https://www.findingtheta.com/blog/mastering-robotic-manipulation-with-reinforcement-learning-tqc-and-ddpg-for-fetch-environments/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/mastering-robotic-manipulation-with-reinforcement-learning-tqc-and-ddpg-for-fetch-environments/</guid><description>Using reinforcement learning (RL), specifically Truncated Quantile Critics (TQC) and Deep Deterministic Policy Gradient (DDPG), to solve the Fetch environments in Gymnasium Robotics</description><pubDate>Wed, 01 Jan 2025 00:00:00 GMT</pubDate><category>experiments</category><category>gymnasium</category><category>robotics</category></item><item><title>Beginner&apos;s Guide to Model-Based Reinforcement Learning (MBRL) with Atari&apos;s Breakout</title><link>https://www.findingtheta.com/blog/beginners-guide-to-model-based-reinforcement-learning-mbrl-with-ataris-breakout/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/beginners-guide-to-model-based-reinforcement-learning-mbrl-with-ataris-breakout/</guid><description>A beginner&apos;s walkthrough of Model-Based Reinforcement Learning on Atari&apos;s Breakout: training an action-conditional world model that predicts frames, rewards and terminals, then learning a PPO policy entirely inside it with SimPLe</description><pubDate>Sun, 01 Dec 2024 00:00:00 GMT</pubDate><category>experiments</category><category>atari</category><category>reinforcement-learning</category></item><item><title>Using Multi-Agent Reinforcement Learning to play OpenSpiel&apos;s Connect 4 with Ray&apos;s RLlib</title><link>https://www.findingtheta.com/blog/using-multi-agent-reinforcement-learning-to-play-openspiels-connect-4-with-rays-rllib/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/using-multi-agent-reinforcement-learning-to-play-openspiels-connect-4-with-rays-rllib/</guid><description>Discover how self-play can be used to train a reinforcement learning agent to master Connect 4, achieving advanced strategies without human intervention</description><pubDate>Fri, 01 Nov 2024 00:00:00 GMT</pubDate><category>experiments</category><category>reinforcement-learning</category></item><item><title>Mastering Atari&apos;s Pong with Reinforcement Learning: Overcoming Sparse Rewards and Optimizing Performance</title><link>https://www.findingtheta.com/blog/mastering-ataris-pong-with-reinforcement-learning-overcoming-sparse-rewards-and-optimizing-performance/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/mastering-ataris-pong-with-reinforcement-learning-overcoming-sparse-rewards-and-optimizing-performance/</guid><description>Train an RL agent to master Atari&apos;s Pong with sparse rewards and high-dimensional inputs. Explore preprocessing, replay buffers, and performance-boosting strategies</description><pubDate>Tue, 01 Oct 2024 00:00:00 GMT</pubDate><category>experiments</category><category>gymnasium</category><category>dqn</category><category>ppo</category><category>atari</category></item><item><title>Solving Gymnasium&apos;s Car Racing with Reinforcement Learning</title><link>https://www.findingtheta.com/blog/solving-gymnasiums-car-racing-with-reinforcement-learning/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/solving-gymnasiums-car-racing-with-reinforcement-learning/</guid><description>Learn how to apply reinforcement learning to solve Gymnasium&apos;s Car Racing game, see how different algorithms perform, and explore whether discrete or continuous action spaces are better.</description><pubDate>Sun, 01 Sep 2024 00:00:00 GMT</pubDate><category>experiments</category><category>gymnasium</category><category>ppo</category><category>dqn</category><category>sac</category></item><item><title>Comparing how PPO, SAC, and DQN Perform on Gymnasium&apos;s Lunar Lander</title><link>https://www.findingtheta.com/blog/comparing-how-ppo-sac-and-dqn-perform-on-gymnasiums-lunar-lander/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/comparing-how-ppo-sac-and-dqn-perform-on-gymnasiums-lunar-lander/</guid><description>Explore how different On-Policy and Off-Policy reinforcement learning algorithms perform on Gymnasium&apos;s Lunar Lander</description><pubDate>Thu, 01 Aug 2024 00:00:00 GMT</pubDate><category>experiments</category><category>dqn</category><category>gymnasium</category><category>ppo</category><category>sac</category></item><item><title>Solving Gymnasium&apos;s Lunar Lander with Deep Q Learning (DQN)</title><link>https://www.findingtheta.com/blog/solving-gymnasiums-lunar-lander-with-deep-q-learning-dqn/</link><guid isPermaLink="true">https://www.findingtheta.com/blog/solving-gymnasiums-lunar-lander-with-deep-q-learning-dqn/</guid><description>Learn how the Reinforcement Learning Algorithm Deep Q Learning (DQN) works and apply it to solve Gymnasium&apos;s Lunar Lander</description><pubDate>Mon, 01 Jul 2024 00:00:00 GMT</pubDate><category>experiments</category><category>dqn</category><category>gymnasium</category><category>reinforcement-learning</category></item></channel></rss>