diff --git a/semester4/dmdb/data-modelling-databases-summary.pdf b/semester4/dmdb/data-modelling-databases-summary.pdf index a2717c7..73c7ed3 100644 Binary files a/semester4/dmdb/data-modelling-databases-summary.pdf and b/semester4/dmdb/data-modelling-databases-summary.pdf differ diff --git a/semester4/dmdb/data-modelling-databases-summary.tex b/semester4/dmdb/data-modelling-databases-summary.tex index cc70759..0caf89b 100644 --- a/semester4/dmdb/data-modelling-databases-summary.tex +++ b/semester4/dmdb/data-modelling-databases-summary.tex @@ -82,6 +82,7 @@ In addition, the end user does not have to worry about optimization too much. Th \include{parts/00_relational-model_logic/main.tex} \include{parts/01_sql/main.tex} \include{parts/02_theory-background/main.tex} +\include{parts/03_systems/main.tex} \printGlossary diff --git a/semester4/dmdb/glossary/defs.tex b/semester4/dmdb/glossary/defs.tex new file mode 100644 index 0000000..241e0c7 --- /dev/null +++ b/semester4/dmdb/glossary/defs.tex @@ -0,0 +1,4 @@ +\newglossaryentry{superkey}{ + name={superkey}, + description={See Section \ref{sec:functional-dependency}} +} diff --git a/semester4/dmdb/glossary/main.tex b/semester4/dmdb/glossary/main.tex index 05e8c02..6a825a9 100644 --- a/semester4/dmdb/glossary/main.tex +++ b/semester4/dmdb/glossary/main.tex @@ -1,3 +1,4 @@ \input{glossary/acronyms.tex} \input{glossary/basics.tex} -% \input{} +\input{glossary/defs.tex} +% \input{glossary/} diff --git a/semester4/dmdb/parts/00_relational-model_logic/00_model/01_keys.tex b/semester4/dmdb/parts/00_relational-model_logic/00_model/01_keys.tex index 4439c87..9856581 100644 --- a/semester4/dmdb/parts/00_relational-model_logic/00_model/01_keys.tex +++ b/semester4/dmdb/parts/00_relational-model_logic/00_model/01_keys.tex @@ -1,4 +1,5 @@ \subsubsection{Keys} +\label{sec:relations-keys} \inlinedefinition[Candidate Key] The minimal set of fields that identify each tuple uniquely \inlinedefinition[Primary Key] A single candidate key, i.e. just a single field. diff --git a/semester4/dmdb/parts/00_relational-model_logic/01_algebra/00_operators.tex b/semester4/dmdb/parts/00_relational-model_logic/01_algebra/00_operators.tex index 3ff231f..bbd7104 100644 --- a/semester4/dmdb/parts/00_relational-model_logic/01_algebra/00_operators.tex +++ b/semester4/dmdb/parts/00_relational-model_logic/01_algebra/00_operators.tex @@ -1,3 +1,4 @@ +\label{sec:relational-algebra} \subsubsection{Operators} \begin{itemize} \item \bi{Union} $\cup$: $x \in R_1 \cup R_2 \Leftrightarrow x \in R_1 \lor x \in R_2$ diff --git a/semester4/dmdb/parts/02_theory-background/00_functional-dependencies/00_intro.tex b/semester4/dmdb/parts/02_theory-background/00_functional-dependencies/00_intro.tex index cfd40b8..eefd990 100644 --- a/semester4/dmdb/parts/02_theory-background/00_functional-dependencies/00_intro.tex +++ b/semester4/dmdb/parts/02_theory-background/00_functional-dependencies/00_intro.tex @@ -1,4 +1,5 @@ \subsection{Functional Dependency} +\label{sec:functional-dependency} As with relational algebra, we again define a schema to be a relation, e.g. $\cR(A:D_A, B:D_B, C:D_C, D: D_D)$, with an instance of it being $R \subseteq D_A \times D_B \times D_C \times D_D$. @@ -12,9 +13,10 @@ and $R \vdash \alpha \rightarrow \beta$ if the same applies syntactically (i.e. \inlineintuition This means that there is a function that maps the values of columns $\alpha$ to the values of columns $\beta$. -\inlinedefinition $\alpha \subseteq \R$ is a \bi{superkey} if and only if $\alpha \rightarrow \cR$. +\inlinedefinition $\alpha \subseteq \R$ is a \gls{superkey} if and only if $\alpha \rightarrow \cR$. -\inlineintuition This means that if we know the values of columns $\alpha$, we know the value of the rest of the columns in $\cR$ +\inlineintuition This means that if we know the values of columns $\alpha$, we know the value of the rest of the columns in $\cR$. +It is a superset of some candidate keys (see \ref{sec:relations-keys}) \inlinedefinition $\alpha \rightarrow \beta$ is minimal if and only if $\forall A \in \alpha : (\alpha - \{ A \}) \centernot{\rightarrow} \beta$. diff --git a/semester4/dmdb/parts/02_theory-background/00_functional-dependencies/04_closure.tex b/semester4/dmdb/parts/02_theory-background/00_functional-dependencies/04_closure.tex index a6cf3b5..c616bac 100644 --- a/semester4/dmdb/parts/02_theory-background/00_functional-dependencies/04_closure.tex +++ b/semester4/dmdb/parts/02_theory-background/00_functional-dependencies/04_closure.tex @@ -26,7 +26,7 @@ We can use that to check the following: \item $F \vdash \alpha \rightarrow \gamma$: Calculate $\alpha^+$ and check $\gamma \in \alpha^+$ \item $F \vdash G$: For each $\alpha \rightarrow \gamma \in G$, check $F \vdash \alpha \rightarrow \gamma$ \item $F$ equivalent to $G$: Check $F \vdash G$ and $G \vdash F$ - \item $K \subseteq \cR$ superkey for $F$: Check $F \vdash K \rightarrow \cR$: + \item $K \subseteq \cR$ \gls{superkey} for $F$: Check $F \vdash K \rightarrow \cR$: \end{itemize} diff --git a/semester4/dmdb/parts/02_theory-background/01_normal-forms/01_first-nf.tex b/semester4/dmdb/parts/02_theory-background/01_normal-forms/01_first-nf.tex index 899b4e6..d4f3c63 100644 --- a/semester4/dmdb/parts/02_theory-background/01_normal-forms/01_first-nf.tex +++ b/semester4/dmdb/parts/02_theory-background/01_normal-forms/01_first-nf.tex @@ -1,5 +1,5 @@ \subsubsection{First Normal Form} -In the first normal form, only atomic domains are allowed. +In the first normal form, only atomic domains are allowed, i.e. keys are not allowed to be arrays. \inlineintuition This makes the concept of a key easier to define, or well defined diff --git a/semester4/dmdb/parts/02_theory-background/01_normal-forms/02_second-nf.tex b/semester4/dmdb/parts/02_theory-background/01_normal-forms/02_second-nf.tex index 40a3a5b..0e20af8 100644 --- a/semester4/dmdb/parts/02_theory-background/01_normal-forms/02_second-nf.tex +++ b/semester4/dmdb/parts/02_theory-background/01_normal-forms/02_second-nf.tex @@ -1,10 +1,16 @@ \subsection{Second Normal Form} -The two dependencies $\{ \text{Legi}, \text{Nr} \} \rightrightarrow \{ \text{Name}, \text{Semester} \}$, -$\{ \text{Legi} \} \rightrightarrow \{ \text{Name}, \text{Semester} \}$ state the following: +% The two dependencies $\{ \text{Legi}, \text{Nr} \} \rightrightarrow \{ \text{Name}, \text{Semester} \}$, +% $\{ \text{Legi} \} \rightrightarrow \{ \text{Name}, \text{Semester} \}$ state the following: +% +% \begin{itemize} +% \item Legi determines the name and semester of a student. +% \item The first map has to happen, even though mapping $\{ \text{Legi}, \text{Nr} \}$ to $\{ \text{Legi} \}$ would be enough. We thus will have redundancy. +% However, this is the exact relationship that the table is meant to capture +% Thus, $\{ \text{Name}, \text{Semester} \}$ should not be in the table at all. +% \end{itemize} -\begin{itemize} - \item Legi determines the name and semester of a student. - \item The first map has to happen, even though mapping $\{ \text{Legi}, \text{Nr} \}$ to $\{ \text{Legi} \}$ would be enough. We thus will have redundancy. - However, this is the exact relationship that the table is meant to capture - Thus, $\{ \text{Name}, \text{Semester} \}$ should not be in the table at all. -\end{itemize} +For a relation to be in 2NF, every non-key attribute needs to minimally dependent on \bi{every} key, i.e. no attribute depends on part of a key. +If relations are not in 2NF, they may experience insert, update and delete anomalies, +which can lead to incorrect, redundant or inconsistent updates of the relations. + +However, some relations can still suffer from update and / or delete anomalies diff --git a/semester4/dmdb/parts/02_theory-background/01_normal-forms/03_third-nf.tex b/semester4/dmdb/parts/02_theory-background/01_normal-forms/03_third-nf.tex index e69de29..bf409e6 100644 --- a/semester4/dmdb/parts/02_theory-background/01_normal-forms/03_third-nf.tex +++ b/semester4/dmdb/parts/02_theory-background/01_normal-forms/03_third-nf.tex @@ -0,0 +1,11 @@ +\subsubsection{Third Normal Form} +A relation $R$ is in 3NF if and only if for all $\alpha \rightarrow B$, at least one of the following conditions holds: +\begin{itemize} + \item $B \in \alpha$ (then $\alpha \rightarrow B$ is a trivial FD) + \item $B$ is an attribute of at least one key + \item $\alpha$ is a \gls{superkey} of $R$ +\end{itemize} +\inlineintuition if $\alpha \rightarrow B$ does not satisfy any of these conditions, $\alpha$ is a concept in its own right, +thus, the 3NF tries to get rid of ``transitive dependencies'' (e.g. $A \rightarrow B, B \rightarrow C$) + +The 3NF can still experience update and delete anomalies. diff --git a/semester4/dmdb/parts/02_theory-background/01_normal-forms/04_bncf.tex b/semester4/dmdb/parts/02_theory-background/01_normal-forms/04_bncf.tex index e69de29..5882a69 100644 --- a/semester4/dmdb/parts/02_theory-background/01_normal-forms/04_bncf.tex +++ b/semester4/dmdb/parts/02_theory-background/01_normal-forms/04_bncf.tex @@ -0,0 +1,13 @@ +\subsubsection{Boyce-Codd Normal Form (BCNF)} +$R$ is in BCNF if and only if for all $\alpha \rightarrow B$ at least one of the following conditions holds: +\begin{itemize} + \item $B \in \alpha$ (thus, $\alpha \rightarrow B$ is a trivial FD) + \item $\alpha$ is a \gls{superkey} of $R$ +\end{itemize} +\inlineintuition In each relation, you only store the same information once. + +The following FD are okay in both 3NF and BCNF because both left hand sides are candidate keys for the table. +\begin{itemize} + \item $\{ \text{PersNr} \} \rightarrow $ Name, Level, Room, City, Street, Canton + \item $\{ \text{Room} \} \rightarrow $ PersNr +\end{itemize} diff --git a/semester4/dmdb/parts/02_theory-background/01_normal-forms/05_decomposition-algorithm.tex b/semester4/dmdb/parts/02_theory-background/01_normal-forms/05_decomposition-algorithm.tex new file mode 100644 index 0000000..4e859cb --- /dev/null +++ b/semester4/dmdb/parts/02_theory-background/01_normal-forms/05_decomposition-algorithm.tex @@ -0,0 +1,25 @@ +\subsubsection{BCNF Decomposition Algorithm} +\begin{algorithm} + \caption{Decomposition algorithm} + \begin{algorithmic}[1] + \Procedure{Decomposition}{$\cR$} + \State \texttt{result} $\gets \{ \cR \}$ + \While{$\exists \cR_i$ in \texttt{result} such that $\cR_i$ is not in BCFN} + \State let $\alpha \rightarrow \beta$ be the evil FD + \State $\cR_i^1 \gets \alpha \cup \beta$ + \State $\cR_i^2 \gets \cR_i - \beta$ + \State $\texttt{result} \gets (\texttt{result} - \cR_i) \cup \{ \cR_i^1, \cR_i^2 \}$ + \EndWhile + \State \Return \texttt{result} + \EndProcedure + \end{algorithmic} +\end{algorithm} +In words, this algorithm takes as input a schema $\cR$ and outputs a list of schemas $\cL = \cR_1, \ldots, \cR_n$, +such that $\cL$ is a lossless decomposition of $\cR$ (i.e. joining them together will perfectly recover the information in $\cR$) +and each of the $\cR_i$ are in BCNF. + +Of note is that this does \bi{not} preserve the functional dependencies and some data redundancies will still remain, +only the ones caused by functional dependencies. + +In addition, the order of picking the evil FD matters, so we use heuristics, which pick the one with the largest right hand side. + diff --git a/semester4/dmdb/parts/02_theory-background/01_normal-forms/06_synthesis-algorightm.tex b/semester4/dmdb/parts/02_theory-background/01_normal-forms/06_synthesis-algorightm.tex new file mode 100644 index 0000000..a680c6b --- /dev/null +++ b/semester4/dmdb/parts/02_theory-background/01_normal-forms/06_synthesis-algorightm.tex @@ -0,0 +1,19 @@ +\subsubsection{3NF Synthesis Algorithm} +\begin{algorithm} + \caption{Synthesis algorithm} + \begin{algorithmic}[1] + \Procedure{Synthesis}{$\cR$} + % TODO: what is the minimal basis? + \State Compute the minimal basis $F_c$ of $F$ + \For{\textbf{all} $\alpha \rightarrow \beta \in F_c$} + \State create $R_{\alpha \cup \beta}(\alpha \cup \beta)$ + \EndFor + \State If none of the above relations contains a superkey, add a relation with a key + \State Eliminate $R_\alpha$ if there exists $R_\alpha'$ such that $\alpha \subseteq \alpha'$ + \EndProcedure + \end{algorithmic} +\end{algorithm} +In words, this algorithm takes as input a schema $\cR$ and outputs a list of schemas $\cL = \cR_1, \ldots, \cR_n$, +such that $\cL$ is a lossless decomposition of $\cR$ and each of the $\cR_i$ are in 3NF. + +This algorithm preserves all functional dependencies, however does not remove redundancies because it is not BCNF. diff --git a/semester4/dmdb/parts/02_theory-background/main.tex b/semester4/dmdb/parts/02_theory-background/main.tex index e294049..a92ffcd 100644 --- a/semester4/dmdb/parts/02_theory-background/main.tex +++ b/semester4/dmdb/parts/02_theory-background/main.tex @@ -14,6 +14,8 @@ \input{parts/02_theory-background/01_normal-forms/02_second-nf.tex} \input{parts/02_theory-background/01_normal-forms/03_third-nf.tex} \input{parts/02_theory-background/01_normal-forms/04_bncf.tex} +\input{parts/02_theory-background/01_normal-forms/05_decomposition-algorithm.tex} +\input{parts/02_theory-background/01_normal-forms/06_synthesis-algorightm.tex} % \input{parts/02_theory-background/01_normal-forms/} % \input{parts/02_theory-background/} diff --git a/semester4/dmdb/parts/03_systems/00_query-optimization/00_intro.tex b/semester4/dmdb/parts/03_systems/00_query-optimization/00_intro.tex new file mode 100644 index 0000000..842583e --- /dev/null +++ b/semester4/dmdb/parts/03_systems/00_query-optimization/00_intro.tex @@ -0,0 +1,8 @@ +A lot of performance can be gained in a database system by executing the query in the correct order. +For that, a query optimizer exists, which uses a quite large amount of metadata to speed up query execution. + +A query optimizer very broadly works as follows: +\begin{enumerate} + \item It searches the space of equivalent execution plans. This space tends to be huge, so efficient exploration is needed. + \item Of all the discovered plans, the best is chosen using a cost model, which describes the most efficient plan +\end{enumerate} diff --git a/semester4/dmdb/parts/03_systems/00_query-optimization/01_search-space/00_basics.tex b/semester4/dmdb/parts/03_systems/00_query-optimization/01_search-space/00_basics.tex new file mode 100644 index 0000000..715acfe --- /dev/null +++ b/semester4/dmdb/parts/03_systems/00_query-optimization/01_search-space/00_basics.tex @@ -0,0 +1,32 @@ +\subsubsection{Search Space} +To create the search space, the optimizer splits a query into a collection of query blocks, +where each query block has one \texttt{SELECT} and \texttt{FROM} clause and \textit{at most} one \texttt{WHERE}, \texttt{GROUP BY} and \texttt{HAVING} clause. + +Query plans are typically drawn up as trees, as they are typically easier to understand that way than as both SQL query or in relational algebra notation. +In addition, the query parser commonly transforms it into a tree-like data structure, so thinking about query plans in this way is a good habit. +For the symbols used here, see section \ref{sec:relational-algebra}. + + +\begin{center} + \begin{forest} + for tree={ + ellipse, + inner sep=1mm, + s sep=10mm, + l sep=3mm + } + [$\Pi_{\texttt{s.PersNr}, \texttt{t.Grade}}$ + [$\bowtie_{\texttt{t.LectureID = l.LectureID}}$ + [$\bowtie_{\texttt{s.PersNr = t.PersNr}}$ + [$\Sigma_{\texttt{s.semester < 5}}$ + [\texttt{Student}] + ] + [\texttt{Test}] + ] + [$\Sigma_{\texttt{l.Name = 'Databases'}}$ + [\texttt{Lecture}] + ] + ] + ] + \end{forest} +\end{center} diff --git a/semester4/dmdb/parts/03_systems/00_query-optimization/01_search-space/01_rewriting-rules.tex b/semester4/dmdb/parts/03_systems/00_query-optimization/01_search-space/01_rewriting-rules.tex new file mode 100644 index 0000000..847b1d1 --- /dev/null +++ b/semester4/dmdb/parts/03_systems/00_query-optimization/01_search-space/01_rewriting-rules.tex @@ -0,0 +1,21 @@ +\paragraph{Rewriting Rules} +Given a logical plan as above, there are multiple ways in which we can construct a physical plan. +We can use different join orders, decide when selection happens and specific implementations of operators. + +The following rules were discussed in the lectures: +\begin{enumerate}[label=\textbf{R\arabic*}] + \item Conjunctive selection operators can be deconstructed into a sequence of individual selections: + \[ + \sigma_{\theta_1 \land \theta_2} = \sigma_{\theta_1}(\sigma_{\theta_2}(E)) + \] + This is useful to split selections to push them earlier in the execution plan (reduces the number of elements to process subsequently) + \item Selection operators are commutative: + \[ + \sigma_{\theta_1}(\sigma_{\theta_2}(E)) = \sigma_{\theta_2}(\sigma_{\theta_1}(E)) + \] + This is useful to allow to first put the one that e.g. has an index. + \item Only the last in a sequence of projections is needed, others can be omitted + \[ + \Pi_{t_1}(\Pi_{t_2}(E)) = \Pi_{t_1}(E) + \] +\end{enumerate} diff --git a/semester4/dmdb/parts/03_systems/main.tex b/semester4/dmdb/parts/03_systems/main.tex new file mode 100644 index 0000000..276a34d --- /dev/null +++ b/semester4/dmdb/parts/03_systems/main.tex @@ -0,0 +1,8 @@ +\newsection +\section{Systems} +\subsection{Query Optimization} +\input{parts/03_systems/00_query-optimization/00_intro.tex} +\input{parts/03_systems/00_query-optimization/01_search-space/00_basics.tex} +\input{parts/03_systems/00_query-optimization/01_search-space/01_rewriting-rules.tex} +% \input{parts/03_systems/00_query-optimization/01_search-space/} +% \input{parts/03_systems/00_query-optimization/}