Initial commit

This commit is contained in:
xixu-me committed 2024-08-20 16:25:10 +08:00
1 parent 1fb7b586eb
commit 4c20ec342f
549 files changed
+805499

No files matched your search

+274
View File
@@ -0,0 +1,274 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "538cb9d2",
"metadata": {
"origin_pos": 0
},
"source": [
"# Beam Search\n",
":label:`sec_beam-search`\n",
"\n",
"In :numref:`sec_seq2seq`, \n",
"we introduced the encoder--decoder architecture,\n",
"and the standard techniques for training them end-to-end. However, when it came to test-time prediction,\n",
"we mentioned only the *greedy* strategy,\n",
"where we select at each time step \n",
"the token given the highest \n",
"predicted probability of coming next, \n",
"until, at some time step, \n",
"we find that we have predicted\n",
"the special end-of-sequence \"<eos>\" token.\n",
"In this section, we will begin \n",
"by formalizing this *greedy search* strategy\n",
"and identifying some problems \n",
"that practitioners tend to run into.\n",
"Subsequently, we compare this strategy\n",
"with two alternatives:\n",
"*exhaustive search* (illustrative but not practical)\n",
"and *beam search* (the standard method in practice).\n",
"\n",
"Let's begin by setting up our mathematical notation,\n",
"borrowing conventions from :numref:`sec_seq2seq`.\n",
"At any time step $t'$, the decoder outputs \n",
"predictions representing the probability \n",
"of each token in the vocabulary \n",
"coming next in the sequence \n",
"(the likely value of $y_{t'+1}$), \n",
"conditioned on the previous tokens\n",
"$y_1, \\ldots, y_{t'}$ and \n",
"the context variable $\\mathbf{c}$,\n",
"produced by the encoder \n",
"to represent the input sequence.\n",
"To quantify computational cost,\n",
"denote by $\\mathcal{Y}$\n",
"the output vocabulary \n",
"(including the special end-of-sequence token \"<eos>\").\n",
"Let's also specify the maximum number of tokens\n",
"of an output sequence as $T'$.\n",
"Our goal is to search for an ideal output from all \n",
"$\\mathcal{O}(\\left|\\mathcal{Y}\\right|^{T'})$\n",
"possible output sequences.\n",
"Note that this slightly overestimates \n",
"the number of distinct outputs \n",
"because there are no subsequent tokens\n",
"once the \"<eos>\" token occurs.\n",
"However, for our purposes, \n",
"this number roughly captures \n",
"the size of the search space.\n",
"\n",
"\n",
"## Greedy Search\n",
"\n",
"Consider the simple *greedy search* strategy from :numref:`sec_seq2seq`.\n",
"Here, at any time step $t'$, \n",
"we simply select the token \n",
"with the highest conditional probability\n",
"from $\\mathcal{Y}$, i.e., \n",
"\n",
"$$y_{t'} = \\operatorname*{argmax}_{y \\in \\mathcal{Y}} P(y \\mid y_1, \\ldots, y_{t'-1}, \\mathbf{c}).$$\n",
"\n",
"Once our model outputs \"<eos>\" \n",
"(or we reach the maximum length $T'$)\n",
"the output sequence is completed.\n",
"\n",
"This strategy might look reasonable, \n",
"and in fact it is not so bad!\n",
"Considering how computationally undemanding it is,\n",
"you'd be hard pressed to get more bang for your buck. \n",
"However, if we put aside efficiency for a minute,\n",
"it might seem more reasonable to search \n",
"for the *most likely sequence*, \n",
"not the sequence of (greedily selected) *most likely tokens*.\n",
"It turns out that these two objects can be quite different. \n",
"The most likely sequence is the one that maximizes the expression\n",
"$\\prod_{t'=1}^{T'} P(y_{t'} \\mid y_1, \\ldots, y_{t'-1}, \\mathbf{c})$.\n",
"In our machine translation example,\n",
"if the decoder truly recovered the probabilities\n",
"of the underlying generative process, \n",
"then this would give us the most likely translation.\n",
"Unfortunately, there is no guarantee \n",
"that greedy search will give us this sequence.\n",
"\n",
"Let's illustrate it with an example.\n",
"Suppose that there are four tokens \n",
"\"A\", \"B\", \"C\", and \"<eos>\" in the output dictionary.\n",
"In :numref:`fig_s2s-prob1`,\n",
"the four numbers under each time step represent\n",
"the conditional probabilities of generating \"A\", \"B\", \"C\", \n",
"and \"<eos>\" respectively, at that time step.\n",
"\n",
"![At each time step, greedy search selects the token with the highest conditional probability.](../img/s2s-prob1.svg)\n",
":label:`fig_s2s-prob1`\n",
"\n",
"At each time step, greedy search selects \n",
"the token with the highest conditional probability. \n",
"Therefore, the output sequence \"A\", \"B\", \"C\", and \"<eos>\" \n",
"will be predicted (:numref:`fig_s2s-prob1`). \n",
"The conditional probability of this output sequence\n",
"is $0.5\\times0.4\\times0.4\\times0.6 = 0.048$.\n",
"\n",
"\n",
"Next, let's look at another example in :numref:`fig_s2s-prob2`. \n",
"Unlike in :numref:`fig_s2s-prob1`, \n",
"at time step 2 we select the token \"C\", \n",
"which has the *second* highest conditional probability.\n",
"\n",
"![The four numbers under each time step represent \n",
"the conditional probabilities of generating \"A\", \"B\", \"C\", and \"<eos>\" at that time step. \n",
"At time step 2, the token \"C\", which has the second highest conditional probability, \n",
"is selected.](../img/s2s-prob2.svg)\n",
":label:`fig_s2s-prob2`\n",
"\n",
"Since the output subsequences at time steps 1 and 2, \n",
"on which time step 3 is based, \n",
"have changed from \"A\" and \"B\" in :numref:`fig_s2s-prob1` \n",
"to \"A\" and \"C\" in :numref:`fig_s2s-prob2`, \n",
"the conditional probability of each token \n",
"at time step 3 has also changed in :numref:`fig_s2s-prob2`. \n",
"Suppose that we choose the token \"B\" at time step 3. \n",
"Now time step 4 is conditional on\n",
"the output subsequence at the first three time steps\n",
"\"A\", \"C\", and \"B\", \n",
"which has changed from \"A\", \"B\", and \"C\" in :numref:`fig_s2s-prob1`. \n",
"Therefore, the conditional probability of generating \n",
"each token at time step 4 in :numref:`fig_s2s-prob2` \n",
"is also different from that in :numref:`fig_s2s-prob1`. \n",
"As a result, the conditional probability of the output sequence \n",
"\"A\", \"C\", \"B\", and \"<eos>\" in :numref:`fig_s2s-prob2`\n",
"is $0.5\\times0.3 \\times0.6\\times0.6=0.054$, \n",
"which is greater than that of greedy search in :numref:`fig_s2s-prob1`. \n",
"In this example, the output sequence \"A\", \"B\", \"C\", and \"<eos>\" \n",
"obtained by the greedy search is not optimal.\n",
"\n",
"\n",
"\n",
"\n",
"\n",
"## Exhaustive Search\n",
"\n",
"If the goal is to obtain the most likely sequence, \n",
"we may consider using *exhaustive search*: \n",
"enumerate all the possible output sequences \n",
"with their conditional probabilities,\n",
"and then output the one that scores \n",
"the highest predicted probability.\n",
"\n",
"\n",
"While this would certainly give us what we desire,\n",
"it would come at a prohibitive computational cost \n",
"of $\\mathcal{O}(\\left|\\mathcal{Y}\\right|^{T'})$,\n",
"exponential in the sequence length and with an enormous\n",
"base given by the vocabulary size.\n",
"For example, when $|\\mathcal{Y}|=10000$ and $T'=10$, \n",
"both small numbers when compared with ones in real applications, we will need to evaluate $10000^{10} = 10^{40}$ sequences, which is already beyond the capabilities of any foreseeable computers.\n",
"On the other hand, the computational cost of greedy search is \n",
"$\\mathcal{O}(\\left|\\mathcal{Y}\\right|T')$: \n",
"miraculously cheap but far from optimal.\n",
"For example, when $|\\mathcal{Y}|=10000$ and $T'=10$, \n",
"we only need to evaluate $10000\\times10=10^5$ sequences.\n",
"\n",
"\n",
"## Beam Search\n",
"\n",
"You could view sequence decoding strategies as lying on a spectrum,\n",
"with *beam search* striking a compromise \n",
"between the efficiency of greedy search\n",
"and the optimality of exhaustive search.\n",
"The most straightforward version of beam search \n",
"is characterized by a single hyperparameter,\n",
"the *beam size*, $k$.\n",
"Let's explain this terminology.\n",
"At time step 1, we select the $k$ tokens \n",
"with the highest predicted probabilities.\n",
"Each of them will be the first token of \n",
"$k$ candidate output sequences, respectively.\n",
"At each subsequent time step, \n",
"based on the $k$ candidate output sequences\n",
"at the previous time step,\n",
"we continue to select $k$ candidate output sequences \n",
"with the highest predicted probabilities \n",
"from $k\\left|\\mathcal{Y}\\right|$ possible choices.\n",
"\n",
"![The process of beam search (beam size $=2$; maximum length of an output sequence $=3$). The candidate output sequences are $\\mathit{A}$, $\\mathit{C}$, $\\mathit{AB}$, $\\mathit{CE}$, $\\mathit{ABD}$, and $\\mathit{CED}$.](../img/beam-search.svg)\n",
":label:`fig_beam-search`\n",
"\n",
"\n",
":numref:`fig_beam-search` demonstrates the \n",
"process of beam search with an example. \n",
"Suppose that the output vocabulary\n",
"contains only five elements: \n",
"$\\mathcal{Y} = \\{A, B, C, D, E\\}$, \n",
"where one of them is “<eos>”. \n",
"Let the beam size be two and \n",
"the maximum length of an output sequence be three. \n",
"At time step 1, \n",
"suppose that the tokens with the highest conditional probabilities \n",
"$P(y_1 \\mid \\mathbf{c})$ are $A$ and $C$. \n",
"At time step 2, for all $y_2 \\in \\mathcal{Y},$ \n",
"we compute \n",
"\n",
"$$\\begin{aligned}P(A, y_2 \\mid \\mathbf{c}) = P(A \\mid \\mathbf{c})P(y_2 \\mid A, \\mathbf{c}),\\\\ P(C, y_2 \\mid \\mathbf{c}) = P(C \\mid \\mathbf{c})P(y_2 \\mid C, \\mathbf{c}),\\end{aligned}$$ \n",
"\n",
"and pick the largest two among these ten values, say\n",
"$P(A, B \\mid \\mathbf{c})$ and $P(C, E \\mid \\mathbf{c})$.\n",
"Then at time step 3, for all $y_3 \\in \\mathcal{Y}$, we compute \n",
"\n",
"$$\\begin{aligned}P(A, B, y_3 \\mid \\mathbf{c}) = P(A, B \\mid \\mathbf{c})P(y_3 \\mid A, B, \\mathbf{c}),\\\\P(C, E, y_3 \\mid \\mathbf{c}) = P(C, E \\mid \\mathbf{c})P(y_3 \\mid C, E, \\mathbf{c}),\\end{aligned}$$ \n",
"\n",
"and pick the largest two among these ten values, say \n",
"$P(A, B, D \\mid \\mathbf{c})$ and $P(C, E, D \\mid \\mathbf{c}).$\n",
"As a result, we get six candidates output sequences: \n",
"(i) $A$; (ii) $C$; (iii) $A$, $B$; (iv) $C$, $E$; (v) $A$, $B$, $D$; and (vi) $C$, $E$, $D$. \n",
"\n",
"\n",
"In the end, we obtain the set of final candidate output sequences \n",
"based on these six sequences (e.g., discard portions including and after “<eos>”).\n",
"Then we choose the output sequence which maximizes the following score:\n",
"\n",
"$$ \\frac{1}{L^\\alpha} \\log P(y_1, \\ldots, y_{L}\\mid \\mathbf{c}) = \\frac{1}{L^\\alpha} \\sum_{t'=1}^L \\log P(y_{t'} \\mid y_1, \\ldots, y_{t'-1}, \\mathbf{c});$$\n",
":eqlabel:`eq_beam-search-score`\n",
"\n",
"here $L$ is the length of the final candidate sequence \n",
"and $\\alpha$ is usually set to 0.75. \n",
"Since a longer sequence has more logarithmic terms \n",
"in the summation of :eqref:`eq_beam-search-score`,\n",
"the term $L^\\alpha$ in the denominator penalizes\n",
"long sequences.\n",
"\n",
"The computational cost of beam search is $\\mathcal{O}(k\\left|\\mathcal{Y}\\right|T')$. \n",
"This result is in between that of greedy search and that of exhaustive search.\n",
"Greedy search can be treated as a special case of beam search \n",
"arising when the beam size is set to 1.\n",
"\n",
"\n",
"\n",
"\n",
"## Summary\n",
"\n",
"Sequence searching strategies include \n",
"greedy search, exhaustive search, and beam search.\n",
"Beam search provides a trade-off between accuracy and \n",
"computational cost via the flexible choice of the beam size.\n",
"\n",
"\n",
"## Exercises\n",
"\n",
"1. Can we treat exhaustive search as a special type of beam search? Why or why not?\n",
"1. Apply beam search in the machine translation problem in :numref:`sec_seq2seq`. How does the beam size affect the translation results and the prediction speed?\n",
"1. We used language modeling for generating text following user-provided prefixes in :numref:`sec_rnn-scratch`. Which kind of search strategy does it use? Can you improve it?\n",
"\n",
"[Discussions](https://discuss.d2l.ai/t/338)\n"
]
}
],
"metadata": {
"language_info": {
"name": "python"
},
"required_libs": []
},
"nbformat": 4,
"nbformat_minor": 5
}
+294
View File
@@ -0,0 +1,294 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "05af2ce8",
"metadata": {
"origin_pos": 0
},
"source": [
"# Bidirectional Recurrent Neural Networks\n",
":label:`sec_bi_rnn`\n",
"\n",
"So far, our working example of a sequence learning task has been language modeling,\n",
"where we aim to predict the next token given all previous tokens in a sequence. \n",
"In this scenario, we wish only to condition upon the leftward context,\n",
"and thus the unidirectional chaining of a standard RNN seems appropriate. \n",
"However, there are many other sequence learning tasks contexts \n",
"where it is perfectly fine to condition the prediction at every time step\n",
"on both the leftward and the rightward context. \n",
"Consider, for example, part of speech detection. \n",
"Why shouldn't we take the context in both directions into account\n",
"when assessing the part of speech associated with a given word?\n",
"\n",
"Another common task---often useful as a pretraining exercise\n",
"prior to fine-tuning a model on an actual task of interest---is\n",
"to mask out random tokens in a text document and then to train \n",
"a sequence model to predict the values of the missing tokens.\n",
"Note that depending on what comes after the blank,\n",
"the likely value of the missing token changes dramatically:\n",
"\n",
"* I am `___`.\n",
"* I am `___` hungry.\n",
"* I am `___` hungry, and I can eat half a pig.\n",
"\n",
"In the first sentence \"happy\" seems to be a likely candidate.\n",
"The words \"not\" and \"very\" seem plausible in the second sentence, \n",
"but \"not\" seems incompatible with the third sentences. \n",
"\n",
"\n",
"Fortunately, a simple technique transforms any unidirectional RNN \n",
"into a bidirectional RNN :cite:`Schuster.Paliwal.1997`.\n",
"We simply implement two unidirectional RNN layers\n",
"chained together in opposite directions \n",
"and acting on the same input (:numref:`fig_birnn`).\n",
"For the first RNN layer,\n",
"the first input is $\\mathbf{x}_1$\n",
"and the last input is $\\mathbf{x}_T$,\n",
"but for the second RNN layer, \n",
"the first input is $\\mathbf{x}_T$\n",
"and the last input is $\\mathbf{x}_1$.\n",
"To produce the output of this bidirectional RNN layer,\n",
"we simply concatenate together the corresponding outputs\n",
"of the two underlying unidirectional RNN layers. \n",
"\n",
"\n",
"![Architecture of a bidirectional RNN.](../img/birnn.svg)\n",
":label:`fig_birnn`\n",
"\n",
"\n",
"Formally for any time step $t$,\n",
"we consider a minibatch input $\\mathbf{X}_t \\in \\mathbb{R}^{n \\times d}$ \n",
"(number of examples $=n$; number of inputs in each example $=d$) \n",
"and let the hidden layer activation function be $\\phi$.\n",
"In the bidirectional architecture,\n",
"the forward and backward hidden states for this time step \n",
"are $\\overrightarrow{\\mathbf{H}}_t \\in \\mathbb{R}^{n \\times h}$ \n",
"and $\\overleftarrow{\\mathbf{H}}_t \\in \\mathbb{R}^{n \\times h}$, respectively,\n",
"where $h$ is the number of hidden units.\n",
"The forward and backward hidden state updates are as follows:\n",
"\n",
"\n",
"$$\n",
"\\begin{aligned}\n",
"\\overrightarrow{\\mathbf{H}}_t &= \\phi(\\mathbf{X}_t \\mathbf{W}_{\\textrm{xh}}^{(f)} + \\overrightarrow{\\mathbf{H}}_{t-1} \\mathbf{W}_{\\textrm{hh}}^{(f)} + \\mathbf{b}_\\textrm{h}^{(f)}),\\\\\n",
"\\overleftarrow{\\mathbf{H}}_t &= \\phi(\\mathbf{X}_t \\mathbf{W}_{\\textrm{xh}}^{(b)} + \\overleftarrow{\\mathbf{H}}_{t+1} \\mathbf{W}_{\\textrm{hh}}^{(b)} + \\mathbf{b}_\\textrm{h}^{(b)}),\n",
"\\end{aligned}\n",
"$$\n",
"\n",
"where the weights $\\mathbf{W}_{\\textrm{xh}}^{(f)} \\in \\mathbb{R}^{d \\times h}, \\mathbf{W}_{\\textrm{hh}}^{(f)} \\in \\mathbb{R}^{h \\times h}, \\mathbf{W}_{\\textrm{xh}}^{(b)} \\in \\mathbb{R}^{d \\times h}, \\textrm{ and } \\mathbf{W}_{\\textrm{hh}}^{(b)} \\in \\mathbb{R}^{h \\times h}$, and the biases $\\mathbf{b}_\\textrm{h}^{(f)} \\in \\mathbb{R}^{1 \\times h}$ and $\\mathbf{b}_\\textrm{h}^{(b)} \\in \\mathbb{R}^{1 \\times h}$ are all the model parameters.\n",
"\n",
"Next, we concatenate the forward and backward hidden states\n",
"$\\overrightarrow{\\mathbf{H}}_t$ and $\\overleftarrow{\\mathbf{H}}_t$\n",
"to obtain the hidden state $\\mathbf{H}_t \\in \\mathbb{R}^{n \\times 2h}$ for feeding into the output layer.\n",
"In deep bidirectional RNNs with multiple hidden layers,\n",
"such information is passed on as *input* to the next bidirectional layer. \n",
"Last, the output layer computes the output \n",
"$\\mathbf{O}_t \\in \\mathbb{R}^{n \\times q}$ (number of outputs $=q$):\n",
"\n",
"$$\\mathbf{O}_t = \\mathbf{H}_t \\mathbf{W}_{\\textrm{hq}} + \\mathbf{b}_\\textrm{q}.$$\n",
"\n",
"Here, the weight matrix $\\mathbf{W}_{\\textrm{hq}} \\in \\mathbb{R}^{2h \\times q}$ \n",
"and the bias $\\mathbf{b}_\\textrm{q} \\in \\mathbb{R}^{1 \\times q}$ \n",
"are the model parameters of the output layer. \n",
"While technically, the two directions can have different numbers of hidden units,\n",
"this design choice is seldom made in practice. \n",
"We now demonstrate a simple implementation of a bidirectional RNN.\n"
]
},
{
"cell_type": "code",
"execution_count": 1,
"id": "ee07527f",
"metadata": {
"execution": {
"iopub.execute_input": "2023-08-18T19:39:47.754935Z",
"iopub.status.busy": "2023-08-18T19:39:47.754598Z",
"iopub.status.idle": "2023-08-18T19:39:51.281601Z",
"shell.execute_reply": "2023-08-18T19:39:51.280325Z"
},
"origin_pos": 3,
"tab": [
"pytorch"
]
},
"outputs": [],
"source": [
"import torch\n",
"from torch import nn\n",
"from d2l import torch as d2l"
]
},
{
"cell_type": "markdown",
"id": "3f2aafd5",
"metadata": {
"origin_pos": 6
},
"source": [
"## Implementation from Scratch\n",
"\n",
"To implement a bidirectional RNN from scratch, we can\n",
"include two unidirectional `RNNScratch` instances\n",
"with separate learnable parameters.\n"
]
},
{
"cell_type": "code",
"execution_count": 2,
"id": "3282034f",
"metadata": {
"execution": {
"iopub.execute_input": "2023-08-18T19:39:51.287307Z",
"iopub.status.busy": "2023-08-18T19:39:51.286196Z",
"iopub.status.idle": "2023-08-18T19:39:51.293977Z",
"shell.execute_reply": "2023-08-18T19:39:51.293017Z"
},
"origin_pos": 7,
"tab": [
"pytorch"
]
},
"outputs": [],
"source": [
"class BiRNNScratch(d2l.Module):\n",
" def __init__(self, num_inputs, num_hiddens, sigma=0.01):\n",
" super().__init__()\n",
" self.save_hyperparameters()\n",
" self.f_rnn = d2l.RNNScratch(num_inputs, num_hiddens, sigma)\n",
" self.b_rnn = d2l.RNNScratch(num_inputs, num_hiddens, sigma)\n",
" self.num_hiddens *= 2 # The output dimension will be doubled"
]
},
{
"cell_type": "markdown",
"id": "01d59627",
"metadata": {
"origin_pos": 9
},
"source": [
"States of forward and backward RNNs\n",
"are updated separately,\n",
"while outputs of these two RNNs are concatenated.\n"
]
},
{
"cell_type": "code",
"execution_count": 3,
"id": "bf749cb8",
"metadata": {
"execution": {
"iopub.execute_input": "2023-08-18T19:39:51.298781Z",
"iopub.status.busy": "2023-08-18T19:39:51.297847Z",
"iopub.status.idle": "2023-08-18T19:39:51.305149Z",
"shell.execute_reply": "2023-08-18T19:39:51.304220Z"
},
"origin_pos": 10,
"tab": [
"pytorch"
]
},
"outputs": [],
"source": [
"@d2l.add_to_class(BiRNNScratch)\n",
"def forward(self, inputs, Hs=None):\n",
" f_H, b_H = Hs if Hs is not None else (None, None)\n",
" f_outputs, f_H = self.f_rnn(inputs, f_H)\n",
" b_outputs, b_H = self.b_rnn(reversed(inputs), b_H)\n",
" outputs = [torch.cat((f, b), -1) for f, b in zip(\n",
" f_outputs, reversed(b_outputs))]\n",
" return outputs, (f_H, b_H)"
]
},
{
"cell_type": "markdown",
"id": "ecf0384a",
"metadata": {
"origin_pos": 11
},
"source": [
"## Concise Implementation\n"
]
},
{
"cell_type": "markdown",
"id": "73c382cc",
"metadata": {
"origin_pos": 12,
"tab": [
"pytorch"
]
},
"source": [
"Using the high-level APIs,\n",
"we can implement bidirectional RNNs more concisely.\n",
"Here we take a GRU model as an example.\n"
]
},
{
"cell_type": "code",
"execution_count": 4,
"id": "e70cf71c",
"metadata": {
"execution": {
"iopub.execute_input": "2023-08-18T19:39:51.309028Z",
"iopub.status.busy": "2023-08-18T19:39:51.308375Z",
"iopub.status.idle": "2023-08-18T19:39:51.313930Z",
"shell.execute_reply": "2023-08-18T19:39:51.312960Z"
},
"origin_pos": 14,
"tab": [
"pytorch"
]
},
"outputs": [],
"source": [
"class BiGRU(d2l.RNN):\n",
" def __init__(self, num_inputs, num_hiddens):\n",
" d2l.Module.__init__(self)\n",
" self.save_hyperparameters()\n",
" self.rnn = nn.GRU(num_inputs, num_hiddens, bidirectional=True)\n",
" self.num_hiddens *= 2"
]
},
{
"cell_type": "markdown",
"id": "9b98765b",
"metadata": {
"origin_pos": 15
},
"source": [
"## Summary\n",
"\n",
"In bidirectional RNNs, the hidden state for each time step is simultaneously determined by the data prior to and after the current time step. Bidirectional RNNs are mostly useful for sequence encoding and the estimation of observations given bidirectional context. Bidirectional RNNs are very costly to train due to long gradient chains.\n",
"\n",
"## Exercises\n",
"\n",
"1. If the different directions use a different number of hidden units, how will the shape of $\\mathbf{H}_t$ change?\n",
"1. Design a bidirectional RNN with multiple hidden layers.\n",
"1. Polysemy is common in natural languages. For example, the word \"bank\" has different meanings in contexts “i went to the bank to deposit cash” and “i went to the bank to sit down”. How can we design a neural network model such that given a context sequence and a word, a vector representation of the word in the correct context will be returned? What type of neural architectures is preferred for handling polysemy?\n"
]
},
{
"cell_type": "markdown",
"id": "c5aab116",
"metadata": {
"origin_pos": 17,
"tab": [
"pytorch"
]
},
"source": [
"[Discussions](https://discuss.d2l.ai/t/1059)\n"
]
}
],
"metadata": {
"language_info": {
"name": "python"
},
"required_libs": []
},
"nbformat": 4,
"nbformat_minor": 5
}
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,271 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "d3d85837",
"metadata": {
"origin_pos": 1
},
"source": [
"# The Encoder--Decoder Architecture\n",
":label:`sec_encoder-decoder`\n",
"\n",
"In general sequence-to-sequence problems\n",
"like machine translation\n",
"(:numref:`sec_machine_translation`),\n",
"inputs and outputs are of varying lengths\n",
"that are unaligned.\n",
"The standard approach to handling this sort of data\n",
"is to design an *encoder--decoder* architecture (:numref:`fig_encoder_decoder`)\n",
"consisting of two major components:\n",
"an *encoder* that takes a variable-length sequence as input,\n",
"and a *decoder* that acts as a conditional language model,\n",
"taking in the encoded input\n",
"and the leftwards context of the target sequence\n",
"and predicting the subsequent token in the target sequence.\n",
"\n",
"\n",
"![The encoder--decoder architecture.](../img/encoder-decoder.svg)\n",
":label:`fig_encoder_decoder`\n",
"\n",
"Let's take machine translation from English to French as an example.\n",
"Given an input sequence in English:\n",
"\"They\", \"are\", \"watching\", \".\",\n",
"this encoder--decoder architecture\n",
"first encodes the variable-length input into a state,\n",
"then decodes the state\n",
"to generate the translated sequence,\n",
"token by token, as output:\n",
"\"Ils\", \"regardent\", \".\".\n",
"Since the encoder--decoder architecture\n",
"forms the basis of different sequence-to-sequence models\n",
"in subsequent sections,\n",
"this section will convert this architecture\n",
"into an interface that will be implemented later.\n"
]
},
{
"cell_type": "code",
"execution_count": 1,
"id": "f6ad17a2",
"metadata": {
"execution": {
"iopub.execute_input": "2023-08-18T19:29:03.809994Z",
"iopub.status.busy": "2023-08-18T19:29:03.809520Z",
"iopub.status.idle": "2023-08-18T19:29:06.566331Z",
"shell.execute_reply": "2023-08-18T19:29:06.565200Z"
},
"origin_pos": 3,
"tab": [
"pytorch"
]
},
"outputs": [],
"source": [
"from torch import nn\n",
"from d2l import torch as d2l"
]
},
{
"cell_type": "markdown",
"id": "15bee8b2",
"metadata": {
"origin_pos": 6
},
"source": [
"## (**Encoder**)\n",
"\n",
"In the encoder interface,\n",
"we just specify that\n",
"the encoder takes variable-length sequences as input `X`.\n",
"The implementation will be provided\n",
"by any model that inherits this base `Encoder` class.\n"
]
},
{
"cell_type": "code",
"execution_count": 2,
"id": "bb981983",
"metadata": {
"execution": {
"iopub.execute_input": "2023-08-18T19:29:06.571309Z",
"iopub.status.busy": "2023-08-18T19:29:06.570433Z",
"iopub.status.idle": "2023-08-18T19:29:06.577664Z",
"shell.execute_reply": "2023-08-18T19:29:06.576626Z"
},
"origin_pos": 8,
"tab": [
"pytorch"
]
},
"outputs": [],
"source": [
"class Encoder(nn.Module): #@save\n",
" \"\"\"The base encoder interface for the encoder--decoder architecture.\"\"\"\n",
" def __init__(self):\n",
" super().__init__()\n",
"\n",
" # Later there can be additional arguments (e.g., length excluding padding)\n",
" def forward(self, X, *args):\n",
" raise NotImplementedError"
]
},
{
"cell_type": "markdown",
"id": "01ce15dd",
"metadata": {
"origin_pos": 11
},
"source": [
"## [**Decoder**]\n",
"\n",
"In the following decoder interface,\n",
"we add an additional `init_state` method\n",
"to convert the encoder output (`enc_all_outputs`)\n",
"into the encoded state.\n",
"Note that this step\n",
"may require extra inputs,\n",
"such as the valid length of the input,\n",
"which was explained\n",
"in :numref:`sec_machine_translation`.\n",
"To generate a variable-length sequence token by token,\n",
"every time the decoder may map an input\n",
"(e.g., the generated token at the previous time step)\n",
"and the encoded state\n",
"into an output token at the current time step.\n"
]
},
{
"cell_type": "code",
"execution_count": 3,
"id": "7874f0f6",
"metadata": {
"execution": {
"iopub.execute_input": "2023-08-18T19:29:06.581746Z",
"iopub.status.busy": "2023-08-18T19:29:06.581040Z",
"iopub.status.idle": "2023-08-18T19:29:06.587735Z",
"shell.execute_reply": "2023-08-18T19:29:06.586641Z"
},
"origin_pos": 13,
"tab": [
"pytorch"
]
},
"outputs": [],
"source": [
"class Decoder(nn.Module): #@save\n",
" \"\"\"The base decoder interface for the encoder--decoder architecture.\"\"\"\n",
" def __init__(self):\n",
" super().__init__()\n",
"\n",
" # Later there can be additional arguments (e.g., length excluding padding)\n",
" def init_state(self, enc_all_outputs, *args):\n",
" raise NotImplementedError\n",
"\n",
" def forward(self, X, state):\n",
" raise NotImplementedError"
]
},
{
"cell_type": "markdown",
"id": "ddb7e5a2",
"metadata": {
"origin_pos": 16
},
"source": [
"## [**Putting the Encoder and Decoder Together**]\n",
"\n",
"In the forward propagation,\n",
"the output of the encoder\n",
"is used to produce the encoded state,\n",
"and this state will be further used\n",
"by the decoder as one of its input.\n"
]
},
{
"cell_type": "code",
"execution_count": 4,
"id": "a4223135",
"metadata": {
"execution": {
"iopub.execute_input": "2023-08-18T19:29:06.591527Z",
"iopub.status.busy": "2023-08-18T19:29:06.590891Z",
"iopub.status.idle": "2023-08-18T19:29:06.597930Z",
"shell.execute_reply": "2023-08-18T19:29:06.596992Z"
},
"origin_pos": 17,
"tab": [
"pytorch"
]
},
"outputs": [],
"source": [
"class EncoderDecoder(d2l.Classifier): #@save\n",
" \"\"\"The base class for the encoder--decoder architecture.\"\"\"\n",
" def __init__(self, encoder, decoder):\n",
" super().__init__()\n",
" self.encoder = encoder\n",
" self.decoder = decoder\n",
"\n",
" def forward(self, enc_X, dec_X, *args):\n",
" enc_all_outputs = self.encoder(enc_X, *args)\n",
" dec_state = self.decoder.init_state(enc_all_outputs, *args)\n",
" # Return decoder output only\n",
" return self.decoder(dec_X, dec_state)[0]"
]
},
{
"cell_type": "markdown",
"id": "b8f923bd",
"metadata": {
"origin_pos": 20
},
"source": [
"In the next section,\n",
"we will see how to apply RNNs to design\n",
"sequence-to-sequence models based on\n",
"this encoder--decoder architecture.\n",
"\n",
"\n",
"## Summary\n",
"\n",
"Encoder-decoder architectures\n",
"can handle inputs and outputs\n",
"that both consist of variable-length sequences\n",
"and thus are suitable for sequence-to-sequence problems\n",
"such as machine translation.\n",
"The encoder takes a variable-length sequence as input\n",
"and transforms it into a state with a fixed shape.\n",
"The decoder maps the encoded state of a fixed shape\n",
"to a variable-length sequence.\n",
"\n",
"\n",
"## Exercises\n",
"\n",
"1. Suppose that we use neural networks to implement the encoder--decoder architecture. Do the encoder and the decoder have to be the same type of neural network?\n",
"1. Besides machine translation, can you think of another application where the encoder--decoder architecture can be applied?\n"
]
},
{
"cell_type": "markdown",
"id": "ae1dc5a6",
"metadata": {
"origin_pos": 22,
"tab": [
"pytorch"
]
},
"source": [
"[Discussions](https://discuss.d2l.ai/t/1061)\n"
]
}
],
"metadata": {
"language_info": {
"name": "python"
},
"required_libs": []
},
"nbformat": 4,
"nbformat_minor": 5
}
File diff suppressed because it is too large. Load diff
+96
View File
@@ -0,0 +1,96 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "2b57309f",
"metadata": {
"origin_pos": 0
},
"source": [
"# Modern Recurrent Neural Networks\n",
":label:`chap_modern_rnn`\n",
"\n",
"The previous chapter introduced the key ideas \n",
"behind recurrent neural networks (RNNs). \n",
"However, just as with convolutional neural networks,\n",
"there has been a tremendous amount of innovation\n",
"in RNN architectures, culminating in several complex\n",
"designs that have proven successful in practice. \n",
"In particular, the most popular designs \n",
"feature mechanisms for mitigating the notorious\n",
"numerical instability faced by RNNs,\n",
"as typified by vanishing and exploding gradients.\n",
"Recall that in :numref:`chap_rnn` we dealt \n",
"with exploding gradients by applying a blunt\n",
"gradient clipping heuristic. \n",
"Despite the efficacy of this hack,\n",
"it leaves open the problem of vanishing gradients. \n",
"\n",
"In this chapter, we introduce the key ideas behind \n",
"the most successful RNN architectures for sequences,\n",
"which stem from two papers.\n",
"The first, *Long Short-Term Memory* :cite:`Hochreiter.Schmidhuber.1997`,\n",
"introduces the *memory cell*, a unit of computation that replaces \n",
"traditional nodes in the hidden layer of a network.\n",
"With these memory cells, networks are able \n",
"to overcome difficulties with training \n",
"encountered by earlier recurrent networks.\n",
"Intuitively, the memory cell avoids \n",
"the vanishing gradient problem\n",
"by keeping values in each memory cell's internal state\n",
"cascading along a recurrent edge with weight 1 \n",
"across many successive time steps. \n",
"A set of multiplicative gates help the network\n",
"to determine not only the inputs to allow \n",
"into the memory state, \n",
"but when the content of the memory state \n",
"should influence the model's output. \n",
"\n",
"The second paper, *Bidirectional Recurrent Neural Networks* :cite:`Schuster.Paliwal.1997`,\n",
"introduces an architecture in which information \n",
"from both the future (subsequent time steps) \n",
"and the past (preceding time steps)\n",
"are used to determine the output \n",
"at any point in the sequence.\n",
"This is in contrast to previous networks, \n",
"in which only past input can affect the output.\n",
"Bidirectional RNNs have become a mainstay \n",
"for sequence labeling tasks in natural language processing,\n",
"among a myriad of other tasks. \n",
"Fortunately, the two innovations are not mutually exclusive, \n",
"and have been successfully combined for phoneme classification\n",
":cite:`Graves.Schmidhuber.2005` and handwriting recognition :cite:`graves2008novel`.\n",
"\n",
"\n",
"The first sections in this chapter will explain the LSTM architecture,\n",
"a lighter-weight version called the gated recurrent unit (GRU),\n",
"the key ideas behind bidirectional RNNs \n",
"and a brief explanation of how RNN layers \n",
"are stacked together to form deep RNNs. \n",
"Subsequently, we will explore the application of RNNs\n",
"in sequence-to-sequence tasks, \n",
"introducing machine translation\n",
"along with key ideas such as *encoder--decoder* architectures and *beam search*.\n",
"\n",
":begin_tab:toc\n",
" - [lstm](lstm.ipynb)\n",
" - [gru](gru.ipynb)\n",
" - [deep-rnn](deep-rnn.ipynb)\n",
" - [bi-rnn](bi-rnn.ipynb)\n",
" - [machine-translation-and-dataset](machine-translation-and-dataset.ipynb)\n",
" - [encoder-decoder](encoder-decoder.ipynb)\n",
" - [seq2seq](seq2seq.ipynb)\n",
" - [beam-search](beam-search.ipynb)\n",
":end_tab:\n"
]
}
],
"metadata": {
"language_info": {
"name": "python"
},
"required_libs": []
},
"nbformat": 4,
"nbformat_minor": 5
}
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff