From Q-Tables to Deep Q-Networks: Scaling Reinforcement Learning with Neural Networks

Author: Yongfeng Gu, Date: October 8, 2026

Q-Learning works beautifully for small problems. But what happens when the state space explodes? In this post, we replace the Q-Table with a neural network and unlock the power of Deep Reinforcement Learning.

The Scalability Problem: When Q-Tables Break

Our 6x6 grid game has 36 states and 4 actions, so its Q-Table contains only 36 x 4 = 144 entries. This is trivially small. A 100x100 grid already has 10,000 states; adding items, battery levels, or moving obstacles makes the state space grow combinatorially. The introduction to the 6x6 grid game please refers to my prior post: Demystifying Q-Learning.

Besides, Real-world problems are even harder. The state space can be effectively infinite, so a Q-Table cannot be stored or learned. This is what we called the curse of dimensionality.

The Core Idea: Neural Networks as Q-Function Approximators

Deep Q-Networks (DQN), introduced by Mnih et al. at DeepMind, replace the Q-Table with a neural network. The network accepts a state as input and outputs Q-values for every action.

Q-Table:  Q(s, a) → lookup table entry
DQN:      Q(s, a; θ) → neural network output

Here, θ denotes the network weights. Rather than updating one table cell at a time, DQN uses gradient descent to reduce the difference between predicted Q-values and target Q-values. A neural network is a generalizer: it can estimate sensible values for a new state from similar states encountered during training.


The 6x6 Grid Game Revisited: Network Architecture

For the grid robot, the network must accept the robot position and output values for Up, Down, Left, and Right. One simple state representation is a 36-dimensional one-hot vector.

Position (2, 3) → index = 2 × 6 + 3 = 15
[0, 0, 0, 0, 0, 0,
 0, 0, 0, 0, 0, 0,
 0, 0, 0, 1, 0, 0,  ← position 15
 0, 0, 0, 0, 0, 0,
 0, 0, 0, 0, 0, 0,
 0, 0, 0, 0, 0, 0]

A fully connected DQN can use two hidden layers and one output layer:

Input Layer:    36 neurons (one-hot state)
     ↓
Hidden Layer 1: 128 neurons, ReLU
     ↓
Hidden Layer 2: 128 neurons, ReLU
     ↓
Output Layer:   4 neurons (Q-values for actions)
import torch.nn as nn

class QNetwork(nn.Module):
    def __init__(self, state_dim=36, action_dim=4):
        super().__init__()
        self.network = nn.Sequential(
            nn.Linear(state_dim, 128), nn.ReLU(),
            nn.Linear(128, 128), nn.ReLU(),
            nn.Linear(128, action_dim))

    def forward(self, state):
        return self.network(state)

For state (2, 3), an output such as [45.2, 78.6, 32.1, 85.3] means Right is currently the best action.


The Training Challenge: Why Naive DQN Fails

Simply training a neural Q-function on sequential experience is unstable. First, consecutive transitions are highly correlated, while neural networks learn best from approximately independent samples. A long sequence in one region can cause overfitting and catastrophic forgetting.

Second, the Bellman target moves as the same network changes:

target = r + γ · max₍a′₎ Q(s′, a′)

Updating network weights changes both the prediction and its target, like chasing a moving target. DQN stabilizes this process with two key innovations.


Innovation 1: Experience Replay

Each transition (s, a, r, s′, done) is stored in a large replay buffer. Training then samples random mini-batches rather than using the newest consecutive transitions.

Replay Buffer
(s₁, a₁, r₁, s₂)  ← oldest
(s₂, a₂, r₂, s₃)
...
(s₉₉₉₉₉, a₉₉₉₉₉, r₉₉₉₉₉, s₁₀₀₀₀₀)  ← newest

Training: randomly sample 32 experiences
  • Breaks correlation by mixing states, episodes, and time steps.
  • Reuses past experience, including rare but valuable failures.
  • Stabilizes learning by presenting a more representative state distribution.
class ReplayBuffer:
    def __init__(self, capacity=100000):
        self.buffer, self.capacity, self.position = [], capacity, 0

    def push(self, experience):
        if len(self.buffer) < self.capacity:
            self.buffer.append(None)
        self.buffer[self.position] = experience
        self.position = (self.position + 1) % self.capacity

    def sample(self, batch_size=32):
        return random.sample(self.buffer, batch_size)

Innovation 2: Target Network

DQN maintains two networks: Q_online, which is trained and selects actions, and Q_target, a frozen copy used only to calculate target values. Periodically, the online weights are copied to the target network.

Q_online (trained every step)       Q_target (frozen)
Weights: θ                    copy → Weights: θ⁻
Action selection                    Target-value computation

The training target becomes:

target = r + γ · max₍a′₎ Q_target(s′, a′; θ⁻)

Because θ⁻ is fixed between updates, the online network can learn toward a stable target. This is the central mechanism that makes DQN training practical.


argmax and max: Action Selection vs. Value Computation

argmax returns the action index, so it answers “which action is best?” max returns the value itself, so it answers “what is the best future value?”

q_values_online = Q_online(state)
best_action = torch.argmax(q_values_online).item()  # e.g. 3 (Right)

q_values_target = Q_target(next_state)
best_q_value = torch.max(q_values_target).item()    # e.g. 84.6

In standard DQN, the online network is used to choose an action while the target network evaluates the next state. Separating these roles helps avoid instability from self-referential targets.


The Complete DQN Algorithm
Initialize replay buffer D
Initialize Q_online with random weights θ
Initialize Q_target with θ⁻ = θ

for each episode:
    reset environment and encode state s

    while episode is not done:
        choose action a using ε-greedy Q_online
        execute a; observe r, next state s′, and done
        store (s, a, r, s′, done) in D

        if D has enough experiences:
            sample a random mini-batch from D
            q_sa = Q_online(states).gather(actions)
            targets = rewards + γ × max(Q_target(next_states)) × (1 - dones)
            minimize MSE(q_sa, targets) with gradient descent

        s ← s′

    periodically copy Q_online weights to Q_target
    decay ε

DQN vs. Q-Learning: A Side-by-Side Comparison
AspectQ-Learning (Tabular)DQN
Q-functionQ-TableNeural network Q(s, a; θ)
State spaceSmall and discreteLarge or continuous
SamplesCurrent sequential episodeRandom replay mini-batches
TargetSame Q-TableSeparate target network
GeneralizationNoneSimilar states share estimates
ConvergenceGuaranteed under assumptionsNo general guarantee, effective in practice

Key insight: DQN retains Q-Learning's Bellman-guided trial-and-error idea. It replaces the lookup table with a neural network and adds replay plus a target network to make large-scale learning stable.


Why Does This Matter for Ad Bidding?

Auto-bidding combines continuous budget, elapsed time, CTR, impressions, competitor behavior, and user features. Even a discretized version has billions of possible states. DQN can learn a Q-function that generalizes: a pattern such as “high budget + early time + high CTR” can inform actions in unseen but similar states.

Modern bidding systems often build on DQN variants such as Double DQN, Dueling DQN, and Rainbow. The core principle remains: learn action-values through experience, then scale the representation with neural networks.


Summary
ConceptWhat It Means
Curse of DimensionalityQ-Tables fail for large or continuous state spaces.
DQNA neural network approximates Q(s, a; θ).
Experience ReplayRandom buffer samples break correlation and reuse experience.
Target NetworkA frozen Q_online copy provides stable targets.
Q_online / Q_targetThe training network / the frozen target-value network.
GeneralizationSimilar states can receive related value estimates.

DQN is a milestone in Deep Reinforcement Learning. With experience replay and target networks, neural Q-functions can solve problems that are impossible for tabular methods. The robot still learns one transition at a time—but it now has a memory of past mistakes, a stable target to pursue, and the capacity to generalize to much larger worlds.