From Q-Tables to Deep Q-Networks: Scaling Reinforcement Learning with Neural Networks
Author: Yongfeng Gu, Date: October 8, 2026
Q-Learning works beautifully for small problems. But what happens when the state space explodes? In this post, we replace the Q-Table with a neural network and unlock the power of Deep Reinforcement Learning.
The Scalability Problem: When Q-Tables Break
Our 6x6 grid game has 36 states and 4 actions, so its Q-Table contains only 36 x 4 = 144 entries. This is trivially small. A 100x100 grid already has 10,000 states; adding items, battery levels, or moving obstacles makes the state space grow combinatorially. The introduction to the 6x6 grid game please refers to my prior post: Demystifying Q-Learning.
Besides, Real-world problems are even harder. The state space can be effectively infinite, so a Q-Table cannot be stored or learned. This is what we called the curse of dimensionality.
The Core Idea: Neural Networks as Q-Function Approximators
Deep Q-Networks (DQN), introduced by Mnih et al. at DeepMind, replace the Q-Table with a neural network. The network accepts a state as input and outputs Q-values for every action.
Q-Table: Q(s, a) → lookup table entry DQN: Q(s, a; θ) → neural network output
Here, θ denotes the network weights. Rather than updating one table cell at a time, DQN uses gradient descent to reduce the difference between predicted Q-values and target Q-values. A neural network is a generalizer: it can estimate sensible values for a new state from similar states encountered during training.
The 6x6 Grid Game Revisited: Network Architecture
For the grid robot, the network must accept the robot position and output values for Up, Down, Left, and Right. One simple state representation is a 36-dimensional one-hot vector.
Position (2, 3) → index = 2 × 6 + 3 = 15 [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, ← position 15 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
A fully connected DQN can use two hidden layers and one output layer:
Input Layer: 36 neurons (one-hot state)
↓
Hidden Layer 1: 128 neurons, ReLU
↓
Hidden Layer 2: 128 neurons, ReLU
↓
Output Layer: 4 neurons (Q-values for actions)
import torch.nn as nn
class QNetwork(nn.Module):
def __init__(self, state_dim=36, action_dim=4):
super().__init__()
self.network = nn.Sequential(
nn.Linear(state_dim, 128), nn.ReLU(),
nn.Linear(128, 128), nn.ReLU(),
nn.Linear(128, action_dim))
def forward(self, state):
return self.network(state)
For state (2, 3), an output such as [45.2, 78.6, 32.1, 85.3] means Right is currently the best action.
The Training Challenge: Why Naive DQN Fails
Simply training a neural Q-function on sequential experience is unstable. First, consecutive transitions are highly correlated, while neural networks learn best from approximately independent samples. A long sequence in one region can cause overfitting and catastrophic forgetting.
Second, the Bellman target moves as the same network changes:
target = r + γ · max₍a′₎ Q(s′, a′)
Updating network weights changes both the prediction and its target, like chasing a moving target. DQN stabilizes this process with two key innovations.
Innovation 1: Experience Replay
Each transition (s, a, r, s′, done) is stored in a large replay buffer. Training then samples random mini-batches rather than using the newest consecutive transitions.
Replay Buffer (s₁, a₁, r₁, s₂) ← oldest (s₂, a₂, r₂, s₃) ... (s₉₉₉₉₉, a₉₉₉₉₉, r₉₉₉₉₉, s₁₀₀₀₀₀) ← newest Training: randomly sample 32 experiences
- Breaks correlation by mixing states, episodes, and time steps.
- Reuses past experience, including rare but valuable failures.
- Stabilizes learning by presenting a more representative state distribution.
class ReplayBuffer:
def __init__(self, capacity=100000):
self.buffer, self.capacity, self.position = [], capacity, 0
def push(self, experience):
if len(self.buffer) < self.capacity:
self.buffer.append(None)
self.buffer[self.position] = experience
self.position = (self.position + 1) % self.capacity
def sample(self, batch_size=32):
return random.sample(self.buffer, batch_size)Innovation 2: Target Network
DQN maintains two networks: Q_online, which is trained and selects actions, and Q_target, a frozen copy used only to calculate target values. Periodically, the online weights are copied to the target network.
Q_online (trained every step) Q_target (frozen) Weights: θ copy → Weights: θ⁻ Action selection Target-value computation
The training target becomes:
target = r + γ · max₍a′₎ Q_target(s′, a′; θ⁻)
Because θ⁻ is fixed between updates, the online network can learn toward a stable target. This is the central mechanism that makes DQN training practical.
argmax and max: Action Selection vs. Value Computation
argmax returns the action index, so it answers “which action is best?” max returns the value itself, so it answers “what is the best future value?”
q_values_online = Q_online(state) best_action = torch.argmax(q_values_online).item() # e.g. 3 (Right) q_values_target = Q_target(next_state) best_q_value = torch.max(q_values_target).item() # e.g. 84.6
In standard DQN, the online network is used to choose an action while the target network evaluates the next state. Separating these roles helps avoid instability from self-referential targets.
The Complete DQN Algorithm
Initialize replay buffer D
Initialize Q_online with random weights θ
Initialize Q_target with θ⁻ = θ
for each episode:
reset environment and encode state s
while episode is not done:
choose action a using ε-greedy Q_online
execute a; observe r, next state s′, and done
store (s, a, r, s′, done) in D
if D has enough experiences:
sample a random mini-batch from D
q_sa = Q_online(states).gather(actions)
targets = rewards + γ × max(Q_target(next_states)) × (1 - dones)
minimize MSE(q_sa, targets) with gradient descent
s ← s′
periodically copy Q_online weights to Q_target
decay εDQN vs. Q-Learning: A Side-by-Side Comparison
| Aspect | Q-Learning (Tabular) | DQN |
|---|---|---|
| Q-function | Q-Table | Neural network Q(s, a; θ) |
| State space | Small and discrete | Large or continuous |
| Samples | Current sequential episode | Random replay mini-batches |
| Target | Same Q-Table | Separate target network |
| Generalization | None | Similar states share estimates |
| Convergence | Guaranteed under assumptions | No general guarantee, effective in practice |
Key insight: DQN retains Q-Learning's Bellman-guided trial-and-error idea. It replaces the lookup table with a neural network and adds replay plus a target network to make large-scale learning stable.
Why Does This Matter for Ad Bidding?
Auto-bidding combines continuous budget, elapsed time, CTR, impressions, competitor behavior, and user features. Even a discretized version has billions of possible states. DQN can learn a Q-function that generalizes: a pattern such as “high budget + early time + high CTR” can inform actions in unseen but similar states.
Modern bidding systems often build on DQN variants such as Double DQN, Dueling DQN, and Rainbow. The core principle remains: learn action-values through experience, then scale the representation with neural networks.
Summary
| Concept | What It Means |
|---|---|
| Curse of Dimensionality | Q-Tables fail for large or continuous state spaces. |
| DQN | A neural network approximates Q(s, a; θ). |
| Experience Replay | Random buffer samples break correlation and reuse experience. |
| Target Network | A frozen Q_online copy provides stable targets. |
| Q_online / Q_target | The training network / the frozen target-value network. |
| Generalization | Similar states can receive related value estimates. |
DQN is a milestone in Deep Reinforcement Learning. With experience replay and target networks, neural Q-functions can solve problems that are impossible for tabular methods. The robot still learns one transition at a time—but it now has a memory of past mistakes, a stable target to pursue, and the capacity to generalize to much larger worlds.
