Training-efficient density quantum machine learning

Machine Learning


Density quantum neural networks

To begin, we explicitly define the framework (see Supplementary Material A for a discussion) of density quantum neural networks (density QNNs) as follows:

$$\rho ({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}}):= \mathop{\sum }\limits_{k=1}^{K}{\alpha }_{k}{U}_{k}({{\boldsymbol{\theta }}}_{k})\rho ({\boldsymbol{x}}){U}_{k}^{\dagger }({{\boldsymbol{\theta }}}_{k})$$

(1)

ρ(x) is a data encoded initial state, which is usually assumed to be prepared via a ‘data-loader’ unitary, \(\rho ({\boldsymbol{x}})=\left\vert {\boldsymbol{x}}\right\rangle \left\langle {\boldsymbol{x}}\right\vert ,\left\vert {\boldsymbol{x}}\right\rangle := V({\boldsymbol{x}}){\left\vert 0\right\rangle }^{\otimes n}\), a collection of sub-unitaries \({\{{U}_{k}\}}_{k = 1}^{K}\), and a distribution, \({\{{\alpha }_{k}\}}_{k = 1}^{K}\), which may depend on x.

For now, we treat the density state above as an abstraction and later in the text we will discuss methods to prepare the state practically and actually use the model. The preparation method will have relevance for the different applications and connections to other paradigms. Once we have chosen a state preparation method for equation (1), we must choose particular specifications for the sub-unitaries. In some cases, we may recast efficiently trainable models/frameworks within the density formalism to increase their expressibility. In others, we use the framework to improve the overall inference speed of models. In this work, we assume that the sub-unitary circuit structures, once chosen, are fixed, and the only trainability arises from the parameters, \({\{{{\boldsymbol{\theta }}}_{k}\}}_{k = 1}^{K}\) therein, as well as the coefficients, \({\{{\alpha }_{k}\}}_{k = 1}^{K}\). In other words, we do not incorporate variable structure circuits learned for example via quantum architecture search.

As a generalisation, one may consider adding a data dependence into the sub-unitary coefficients, αα(x), while retaining the distributional requirement for all x, \(\sum _{k}\alpha {({\boldsymbol{x}})}_{k}=1\). This gives us the more general family of density QNN states:

$${\rho }_{{\mathsf{D}}}({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})=\mathop{\sum }\limits_{k=1}^{K}{\alpha }_{k}({\boldsymbol{x}}){U}_{k}({{\boldsymbol{\theta }}}_{k})\rho ({\boldsymbol{x}}){U}_{k}^{\dagger }({{\boldsymbol{\theta }}}_{k})$$

(2)

In the QML world, overly dense or expressive single unitary models are known to have problems related to trainability via barren plateaus28. Density QNNs and related frameworks may be a useful direction to retain highly parameterised models but via a combination of smaller, trainable models. We will demonstrate this through several examples in the remainder of the text. Before doing so, in the next section, we want to appropriately cast density QNNs within the current spectrum of quantum machine learning models.

Connection to other QML frameworks

Before proceeding, we first discuss the connection to other popular QML frameworks. For supervised learning purposes, each term in the density state, \({U}_{k}({{\boldsymbol{\theta }}}_{k})\rho ({\boldsymbol{x}}){U}_{k}^{\dagger }({{\boldsymbol{\theta }}}_{k})\) is expressive enough by itself to capture most basic models in the literature. This is due to the common model definition as \(f({\boldsymbol{\theta }},{\boldsymbol{x}}):= {\rm{Tr}}({\mathcal{O}}U({\boldsymbol{\theta }})\rho ({\boldsymbol{x}}){U}^{\dagger }({\boldsymbol{\theta }}))={\rm{Tr}}({\mathcal{O}}({\boldsymbol{\theta }})\rho ({\boldsymbol{x}}))\) for some observable, \({\mathcal{O}}\), i.e. the overlap between a parameterised Hermitian observable and a data-dependent state. This unifies many paradigms in quantum machine learning literature such as kernel methods29 and data reuploading models via gate teleportation30. Due to the linearity of the quantum mechanics, we can write it also in this form by inserting equation (2) into the function evaluation:

$$\begin{array}{lll}{f}_{{\mathsf{D}}}(\{{\boldsymbol{\theta }},{\boldsymbol{\alpha }}\},{\boldsymbol{x}})&=&{\rm{Tr}}\left({\mathcal{O}}({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})\rho ({\boldsymbol{x}})\right),\\ {\mathcal{O}}({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})&:=&\mathop{\sum }\limits_{k=1}^{K}{\alpha }_{k}({\boldsymbol{x}}){U}_{k}^{\dagger }({{\boldsymbol{\theta }}}_{k}){\mathcal{O}}{U}_{k}({{\boldsymbol{\theta }}}_{k})\end{array}$$

(3)

Removing this data-dependence from the coefficients simply removes the data-dependence from the observable, \({\mathcal{O}}({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})\to {\mathcal{O}}({\boldsymbol{\theta }},{\boldsymbol{\alpha }})\). Finally, selecting the sub-unitaries to be identical and equal to the data loading unitary, with K equal to the size of the training data leads to observable \({\mathcal{O}}({\{{{\boldsymbol{x}}}_{k}\}}_{k = 1}^{M},{\boldsymbol{\alpha }}):= \mathop{\sum }\nolimits_{k=1}^{M}{\alpha }_{k}{U}^{\dagger }({{\boldsymbol{x}}}_{k}){\mathcal{O}}U({{\boldsymbol{x}}}_{k})=\mathop{\sum }\nolimits_{k=1}^{M}{\alpha }_{k}\rho ({{\boldsymbol{x}}}_{k})\). This is an optimal family of models in a kernel method via the representer theorem29.

Next, returning to equation (3) and replacing the data dependence from the state with a parameter dependence, ρ(x) → ρ(θ), we fall within the family of flipped quantum models31. These are a useful model family where the role of data and parameters in the model have been flipped. This insight enables the incorporation of classical shadows32 for, e.g. quantum training and classical deployment of QML models.

Finally, we have the framework of post-variational quantum models33, originating from the classical combinations of quantum states ansatz34. In motivation, these models are perhaps more similar to the ‘implicit30 models such as quantum kernel methods, where the quantum computer is used only for specific fixed, non-trainable, operations (e.g. evaluating inner products for kernels), rather than ‘explicit30 models where trainable parameters reside within unitaries, U(θ). Post-variational models involve optimising coefficients αk,q, which are injected into the model via a linear or non-linear combination of observables, \({\{{{\mathcal{O}}}_{q}\}}_{q = 1}^{Q}\), applied to (non-trainable) unitary transformed states, \({\{{U}_{k}\rho ({\boldsymbol{x}}){U}_{k}^{\dagger }\}}_{k = 1}^{K}\). In the linear case, the output of the model is:

$$\begin{array}{lll}{f}_{{\mathsf{PV}}}({\boldsymbol{\alpha }},{\boldsymbol{x}})&=&\mathop{\sum}\limits _{kq}{\alpha }_{kq}{\rm{Tr}}({{\mathcal{O}}}_{q}{U}_{k}\rho ({\boldsymbol{x}}){U}_{k}^{\dagger })\\\qquad\qquad& =&\mathop{\sum}\limits _{kq}{\alpha }_{kq}{\rm{Tr}}({{\mathcal{O}}}_{kq}\rho ({\boldsymbol{x}})),{{\mathcal{O}}}_{kq}:= {U}_{k}{{\mathcal{O}}}_{q}{U}_{k}^{\dagger }\end{array}$$

(4)

The major benefit of post-variational models is that, similar to quantum kernel methods, the optimisation over a convex combination of parameters outside the circuit is in principle significantly easier than the non-convex optimisation of parameters within the unitaries. However, just like kernel methods, this comes at the limitation of an expensive forward pass through the model, which requires \({\mathcal{O}}(KQ)\) circuits to be evaluated. In the worst case, this should also be exponential in the number of qubits in to enable arbitrary quantum transformations on ρ(x), KQ ≤ 4n 33. To avoid evaluating an exponential number of quantum circuits, it is clearly necessary to employ heuristic strategies or impose symmetries to choose a sufficiently large yet expressive pool of operators \({{\mathcal{O}}}_{kq}\). In light of this, ref. 33 proposes ansatz expansion strategies34 or gradient heuristics to grow the pool of quantum operations. Such techniques may be also incorporated into our proposal, but we leave such investigations to future work.

Preparing density quantum neural networks

As mentioned above, we have not yet described a method to prepare the density QNN state, equation (1). Figure 1 showcases two methods of doing so. For now, we do not assume any specific choice for the sub-unitaries. There are two methods to prepare the density state. The first is via a deterministic circuit which exactly prepares ρ(θ, α, x), and shown in Fig. 1b. We prove the correctness of this circuit in Supplementary Material I.2. The structure of the circuit can be related directly to the corresponding linear combination of unitaries QNN35 which prepares instead the pure state, \(\sum _{k}{\alpha }_{k}{U}_{k}({{\boldsymbol{\theta }}}_{k})\left\vert {\boldsymbol{x}}\right\rangle\), seen in Fig. 1a. Notably, the deterministic density QNN removes the need for ancilla postselection on a specific state, (\(({\left\vert 0\right\rangle }_{{\mathcal{A}}}^{\otimes n})\) in the figure). In other words, while a single forward pass through an LCU QNN will only succeed with some probability p, the deterministic density QNN state preparation succeeds with probability p = 1. While the circuits in Fig. 1a, b are conceptually simple, the controlled operation of the sub-unitaries may be very expensive in practice, which is a necessity without any further assumptions. In Supplementary Material I.2 we discuss certain assumptions on the structure of the sub-unitaries which may simplify the resource requirements of this preparation mechanism, specifically assuming a Hamming-weight preserving structure allows the removal of the generic controlled operation.

Fig. 1: Density quantum neural networks.
figure 1

a Linear combination of unitaries quantum neural networks (LCU QNNs) preparing the state \(\sum _{k}{\alpha }_{k}{U}_{k}({{\boldsymbol{\theta }}}_{k})\left\vert {\boldsymbol{x}}\right\rangle\) via postselection on an ancilla register \({\mathcal{A}}\) which prepares the distribution α. b shows corresponding density quantum neural network, implemented deterministically to prepare the state ρ(θ, α, x). Finally, the instantiation of the density QNN state via randomisation is shown in (c) where sub-unitary, Uk(θk) is only prepared with the probability αk without the need for the multi controlled deep circuits and ancilla qubits. The deterministic density QNN, (b) is required if one wishes to make a true comparison of these networks to the dropout mechanism. From the Mixing lemma, the randomised version, (c), can distil the performance benefits of the more powerful LCU QNN, (a), into very short depth circuits. The probability loaders, \({\mathsf{Load}}\left(\sqrt{{\boldsymbol{\alpha }}}\right)\) are assumed to be unary data loaders which act on K qubits within the register, \({\mathcal{A}}\) and have depth \(\log (K)\)71. One could also use binary Prepare and Select circuits acting on \(\log (K)\) qubits as is more standard in LCU literature. The resulting functions from each network, f(θ, α, x) result from the measurement of an observable, \({\mathcal{O}}\), via \(f({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})={\rm{Tr}}({\mathcal{O}}\sigma ({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}}))\), where σ is the output state from each circuit. V(x) is the n-qubit data loader acting on register, \({\mathcal{B}}\).

The second method uses the distributional property of α to only prepare the density state ρ(θ, α, x) on average, depicted in Fig. 1c. In this form, the forward pass completely removes the need for ancillary qubits, and complicated controlled unitaries.

The effect of this is threefold:

  • A forward pass through the randomised density QNN (Fig. 1c) requires time which is upper bounded by the execution speed of only the most complex unitary, \({U}_{{k}^{* }}\). This is illustrated in Fig. 1c. In the language of post-variational models measuring Q observables on a randomised density QNN has complexity \({\mathcal{O}}(Q)\), a K-fold improvement.

  • Secondly, we will show that the gain in efficiency in moving from the LCU to randomised density QNN does not come at a significant loss in model performance. We prove this, under certain assumptions, using the Hastings-Campbell Mixing lemma from randomised quantum circuit compiling.

  • Thirdly, and related to the first two points, one can view the randomised density QNN as an explicit version of the post-variational (in the sense of ref. 30) framework. This may be an interesting direction to study given the series of hierarchies found by ref. 30 between implicit, explicit and reuploading models.

Gradient extraction for density QNNs

For density QNNs to be performant in practice, they must be efficiently trainable. In other words, it should not be exponentially more difficult to evaluate gradients from such models, compared to the component sub-unitaries. In the following, we describe general statements regarding the gradient extractability from density QNNs. By then choosing the sub-unitaries to themselves be efficiently trainable (in line with a so-called backpropagation scaling, which we will define), the entire model will also be. We formalise this as follows:

Proposition 1

(Gradient scaling for density quantum neural networks) Given a density QNN as in equation (1) composed of K sub-unitaries, \({\mathcal{U}}={\{{U}_{k}({{\boldsymbol{\theta }}}_{k})\}}_{k = 1}^{K}\), implemented with distribution, α = {αk}, an unbiased estimator of the gradients of a loss function, \({\mathcal{L}}\), defined by a Hermitian observable, \({\mathcal{H}}\):

$${\mathcal{L}}({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})={\rm{Tr}}\left({\mathcal{H}}\rho ({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})\right)$$

(5)

can be computed by classically post-processing \(\mathop{\sum }\nolimits_{l=1}^{K}\mathop{\sum }\nolimits_{k=1}^{K}{T}_{\ell k}\) circuits, where Tk is the number of circuits required to compute the gradient of sub-unitary k, U(θk) with respect to the parameters in sub-unitary , θ. Furthermore, these parameters may be shared across the unitaries, \({{\boldsymbol{\theta }}}_{k}={{\boldsymbol{\theta }}}_{k{\prime} }\) for some \(k,k^{\prime}\).

The proof is given in Supplementary Material B.1, but it follows simply from the linearity of the model. Now, there are two sub-cases one can consider. First, if all parameters between sub-unitaries are independent, θkθ, k, . This gives the following corollary, also in Supplementary Material B.1 and illustrated in Fig. 2.

Fig. 2: Illustration of Corollary 1.
figure 2

In the case where no parameters are shared across the sub-unitaries, the gradients of the density model in equation (1) when measured with an observable \({\mathcal{H}}\) simply involves computing gradients for each sub-unitary individually. As a result, the full model introduces an \({\mathcal{O}}(K)\) overhead for gradient extraction. If \(K={\mathcal{O}}(\log (N))\) and each sub-unitary admits a backpropagation scaling for gradient extraction, the density model will also admit a backpropagation scaling.

Corollary 1

Given a density QNN as in equation (1) composed of K sub-unitaries, \({\mathcal{U}}={\{{U}_{k}({{\boldsymbol{\theta }}}_{k})\}}_{k = 1}^{K}\) where the parameters of sub-unitaries are independent, θkθ, k, an unbiased estimator of the gradients of a loss function, \({\mathcal{L}}\), equation (5) can be computed by classically post-processing \(\mathop{\sum }\nolimits_{k=1}^{K}{T}_{k}\) circuits, where Tk is the number of circuits required to compute the gradient of sub-unitary k, U(θk) with respect to the parameters, θk.

The second case is where some (or all) parameters are shared across the sub-unitaries. Taking the extreme example, \({\theta }_{l}^{\,j}={\theta }_{k}^{\,j}=:{\theta }^{\,j}\,\forall k,l\)—i.e. all sub-unitaries from equation (1) have the same number of parameters, which are all identical. In this case, for each sub-unitary, l, we must evaluate all K terms in the sum so at most the number of circuits will increase by a factor of K2—we need to compute every term in the matrix of partial derivatives.

Note that this is the number of circuits required, not the overall sample complexity of the estimate. For example, take the single layer commuting-block circuit (just a commuting-generator circuit) with C mutually commuting generators. Also assume a suitable measurement observable, \({\mathcal{H}}\), such that the resulting gradient observables, \({\{{{\mathcal{O}}}_{c}| {{\mathcal{O}}}_{c}:= [{G}_{c},{\mathcal{H}}]\}}_{k = 1}^{C}\), can be simultaneously diagonalized. To estimate these C gradient observables each to a precision ε (meaning outputting an estimate \({\tilde{o}}_{c}\) such that \(| {\tilde{o}}_{k}-\left\langle \psi \right\vert {{\mathcal{O}}}_{c}\left\vert \psi \right\rangle | \le \varepsilon\) with confidence 1 − δ) requires \({\mathcal{O}}\left({\varepsilon }^{-2}\log \left(\frac{C}{\delta }\right)\right)\) copies of ψ (or equivalently calls to a unitary preparing ψ). It is also possible to incorporate strategies such as shadow tomography32, amplitude estimation36 or quantum gradient algorithms37 to improve the C, δ or ε parameter scalings for more general scenarios, though inevitably at the cost of scaling in the others.

Efficiently trainable density networks

The results of the previous section state that moving to the density framework does not result in an exponential increase in gradient extraction difficulty, unless the number of sub-unitaries is exponential. However, what we really care for is that the models are end-to-end efficiently trainable, meaning that overall their gradients can be computed with a backpropagation scaling. This is the resource scaling which the (classical) backpropagation algorithm obeys, and which we ideally would strive for in quantum models. In the following, we can specialise the derived results to the cases where the component sub-unitaries have efficient gradient extraction protocols. This will render the entire density model also efficiently trainable in this regime.

This ‘backpropagation’ scaling can be defined as follows. Specifically:

Definition 1

(Backpropagation scaling16,38) Given a parameterised function, \(f({\boldsymbol{\theta }}),{\boldsymbol{\theta }}\in {{\mathbb{R}}}^{N}\), with \(f^{\prime} ({\boldsymbol{\theta }})\) being an estimate of the gradient of f with respect to θ up to some accuracy ε. The total computational cost to estimate \(f^{\prime} ({\boldsymbol{\theta }})\) with backpropagation is bounded with:

$${\mathcal{T}}(f^{\prime} ({\boldsymbol{\theta }}))\le {c}_{t}{\mathcal{T}}(f({\boldsymbol{\theta }}))$$

(6)

and

$${\mathcal{M}}(f^{\prime} ({\boldsymbol{\theta }}))\le {c}_{m}{\mathcal{M}}(f({\boldsymbol{\theta }}))$$

(7)

where \({c}_{t},{c}_{m}={\mathcal{O}}(\log (N))\) and \({\mathcal{T}}(g)/{\mathcal{M}}(g)\) is the time/amount of memory required to compute g.

In plain terms, a model which achieves a backpropagation scaling according to Definition 1, particularly for quantum models, implies that it does not take significantly more effort, (in terms of number of qubits, circuit size, or number of circuits) to compute gradients of the model with respect to all parameters, than it does to evaluate the model itself.

One family of circuits which does obey such a scaling are the so-called commuting-block QNNs, defined in ref. 38, and which contain B blocks of unitaries generated by operators which all mutually commute within a block. We discuss the specific circuits in ‘Methods’ but for now we specialise Proposition 1 to these commuting-block unitaries as follows:

Corollary 2

(Gradient scaling for density commuting-block quantum neural networks) Given a density QNN containing k sub-unitaries, each acting on n qubits. Each sub-unitary, k, has a commuting-block structure with Bk blocks. Assume each sub-unitary has different parameters, θkθ, k, . Then an unbiased estimate of the gradient can be estimated by classical post-processing \({\mathcal{O}}(2\sum _{k}{B}_{k}-K)\) circuits on n + 1 qubits.

Proof

This follows immediately from Proposition 1 and Theorem 5 from ref. 38. Here, the gradients of a single B-block commuting block circuit can be computed by post-processing 2B − 1 circuits, where 2 circuits are required per block, with the exception of the final block, which can be treated as a commuting generator circuit and evaluated with a single circuit.

At this stage, we showcase two possibilities when constructing density networks. It should be noted that these in some sense represent extreme cases, and should not be taken as the exclusive possibilities. Ultimately, the successful models will likely exist in the middle group. The first path allows us to increase the trainability of certain QML models in the literature. In Table 1, we show some results if we were to do so for some popular examples. The first step is to dissect commuting-block components from each ‘layer’ of the respective model, then treat these components as sub-unitaries within the density formalism and then apply Corollary 2. In the following sections, we describe this strategy for the models in the table, beginning with the hardware efficient ansatz.

Table 1 Summary of gradient scalings for training density quantum neural networks

Secondly, we may simply use it as a means to increase overall model expressibility—where the component sub-unitaries, Uk(θk) are any generic trainable circuits (which are independent for simplicity). Further assume each Uk(θk) has identical structure with N parameters acting on n qubits and requiring Tn,N gradient circuits each. Then, from Corollary 1, a density model will require KTn,N parameter-shift circuits. In many cases, K will be a constant independent of N, or n, and furthermore this evaluation over the K sub-unitaries can be done in parallel.

In the next section, and in Fig. 3 we illustrate these paths using hardware efficient QNNs. For the examples in the following sections (the other models referenced in Table 1 and others), we demonstrate both of these directions.

Fig. 3: Decomposing a hardware efficient ansatz for a density QNN.
figure 3

D layers of a hardware efficient (HWE) ansatz with entanglement generated by CNOT ladders and trainable parameters in single qubits Rx, Ry, Rz gates. (bottom left) D layers extracted into D sub-unitaries with probabilities, \({\{{\alpha }_{d}\}}_{d = 1}^{D}\) for a density QNN version. Applying the commuting-generator framework to the density version, \({\rho }^{{\mathsf{HWE}}}({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})\), enables parallel gradient evaluation in 2D circuits versus 2nD as required by the pure state version, \(\left\vert {\psi }^{{\mathsf{HWE}}}({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})\right\rangle\). TO illustrate potential differences between sub-unitaries, we arbitrarily reverse CNOT directions in subsequent layers and partially accounting for low circuit depth. (bottom right) Alternatively, we can simply create a more expressive version of the hardware efficient QNN within the density framework by duplicating across K sub-unitaries with probabilities \({\{{\alpha }_{k}\}}_{k = 1}^{K}\) retaining D layers each. In this case, the model requires 2nDK circuits for gradient extraction, but each sub-unitary can have independent parameters learning different features, especially if each contains different entanglement structures.

Hardware efficient quantum neural networks

To illustrate the two possible paths for model construction, we use a toy example (shown in Fig. 3)—the common but much maligned hardware efficient39 quantum neural network. These ‘problem-independent’ ansätze were proposed to keep quantum learning models as close as possible to the restrictions of physical quantum computers, by enforcing specific qubit connectivities and avoiding injecting trainable parameters into complex transformations. These circuits are extremely flexible, but this comes at the cost of being vulnerable to barren plateaus23 and generally difficult to train.

A D layer hardware efficient ansatz on n qubits is usually defined to have 1 parameter per qubit (located in a single qubit Pauli rotation) per layer. The parameter-shift rule with such a model would require 2nD individual circuits to estimate the full gradient vector, each for M measurements shots. Given such a circuit, we can construct a density version with D sub-unitaries and reduce the gradient requirements from 2nD to 2D as the gradients for the single qubit unitaries in each sub-unitary can be evaluated in parallel, using the commuting-generator toolkit from Methods and Corollary 2. This example is relatively trivial as the resulting unitaries are shallow depth (which also likely increases the ease of classical simulability) and training each corresponds only to learning a restricted single qubit measurement basis. In the Fig. 3, we take a variation of the common CNOT-ladder layout—entanglement is generated in each layer by nearest-neighbour CNOT gates. Typically, an identical structure is used in each layer, however in the figure we allow each sub-unitary extracted from each layer to have a varying CNOT control-target directionality and different single qubit rotations in each layer. This is to increase differences between each ‘expert’ (see below) as each sub-unitary can generate different levels of (dis)entanglement. Secondly, as illustrated in Fig. 3 one can also define a density version which is not more trainable than the original version—in this case, we have K depth-D hardware efficient circuits, which according to the parameter shift rule would now require \({\mathcal{O}}(2KDn)\) circuits. However, the density model contains more parameters (K-fold more) than the original single circuit version, and possibly is more expressive as a QML model.

LCU and the mixing lemma

The second feature of the density framework is the relationship to linear combination of unitaries (LCU) quantum machine learning. Above, we discussed two methods of preparing the density state, equation (1) and illustrated in Fig. 1. Now, we will demonstrate how one may translate performance guarantees from families of LCU QNNs (Fig. 1a) to the randomised version of the density QNN (Fig. 1c).

Specifically, we will show that, in at least one restricted learning scenario, we will show that if one can construct and train an LCU QNN (Fig. 1a) which has a better learning performance (in terms of e.g. classification accuracy) than any component unitary, this improved performance can be transferred to a density QNN without performance loss. This transference has an important consequence—due to the minimal requirements of implementing a randomised density QNN (Fig. 1c) on quantum hardware, relative to the LCU QNN, we can implement the more performant model much more cheaply. To do so, we will prove a result using the Hastings-Campbell mixing lemma40,41 from the field compiling of complex unitaries onto sequences of simpler quantum operations.

In this context, we will adapt the Mixing lemma as follows. Assume one trains K sub-unitaries {Uk(θk)} each to be ‘good’ models, in that they each achieve a low prediction error, δ1, to some ground truth function. Next, with the trained sub-unitaries fixed, one learns a linear combination, \(\sum _{k}{\alpha }_{k}{U}_{k}\), with (distributional) coefficients, {αk}, by training only the coefficients. Assume this more powerful QNN model (LCU QNN) achieves a ‘better’ prediction error, δ2 < δ1. However, despite better performance, the LCU QNN is far more expensive to implement than any individual Uk (as can be seen in Fig. 1a). The logic of the Mixing lemma implies that instead of this deep circuit, we may randomise over the unitaries—create a randomised density QNN—and achieve the same error as the LCU QNN but with the same overhead as the most complex Uk. We formalise this as the following:

Lemma 1

(Mixing lemma for supervised learning) Let h(x) be a target ground truth function, prepared via the application of a fixed unitary, V, \(h({\boldsymbol{x}}):= {\rm{Tr}}({\mathcal{O}}V\rho ({\boldsymbol{x}}){V}^{\dagger })\) on a data encoded state, ρ(x) and measured with a fixed observable, \({\mathcal{O}}\). Suppose there exists K unitaries \({\{{U}_{k}({\boldsymbol{\theta }})\}}_{k = 1}^{K}\) such that these each are δ1 good predictive models of h(x):

$${{\mathbb{E}}}_{{\boldsymbol{x}}}| h({\boldsymbol{x}})-{f}_{k}({\boldsymbol{\theta }},{\boldsymbol{x}})| \le {\delta }_{1},\forall k$$

(8)

and a distribution \({\{{\alpha }_{k}\}}_{k = 1}^{K}\) such that predictions according to the LCU model \({f}_{{\mathsf{LCU}}}({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}}):= {\rm{Tr}}({\mathcal{O}}\left(\sum _{k}{\alpha }_{k}{U}_{k}({\boldsymbol{\theta }})\right)\rho ({\boldsymbol{x}})\left(\sum _{k}{\alpha }_{k}{U}_{k}^{\dagger }({\boldsymbol{\theta }})\right))\) have error bounded as:

$${{\mathbb{E}}}_{{\boldsymbol{x}}}| h({\boldsymbol{x}})-{f}_{{\mathsf{LCU}}}({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})| \le {\delta }_{2}$$

(9)

for some δ1, δ2 > 0. Then, the corresponding density QNN, \(f({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})={\rm{Tr}}({\mathcal{O}}\rho ({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})),\rho ({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})=\mathop{\sum }\nolimits_{k=1}^{K}{\alpha }_{k}{U}_{k}^{\dagger }({\boldsymbol{\theta }})\rho ({\boldsymbol{x}}){U}_{k}({\boldsymbol{\theta }})\) can generate predictions for h(x) with error:

$${{\mathbb{E}}}_{{\boldsymbol{x}}}| h({\boldsymbol{x}})-f({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})| \le \frac{{\delta }_{1}^{2}}{4\parallel {\mathcal{O}}{\parallel }_{\infty }}+2{\delta }_{2}$$

(10)

We prove this lemma in Supplementary Material B.2. In the above, \({{\mathbb{E}}}_{{\boldsymbol{x}}}| h({\boldsymbol{x}})-{g}_{{\boldsymbol{\theta }}}({\boldsymbol{x}})|\) is the expected prediction error admitted by the model, gθ(x), where the expectation is taken over the distribution the data is drawn from. Take the common predictor observable to be \({\mathcal{O}}=Z\), \(\parallel {\mathcal{O}}{\parallel }_{\infty }=1\) and set δ1 = δ. Assume we find a distribution for the LCU QNN which quadratically reduces this error δ2 = δ2. Then according to Lemma 1, the density QNN will also be an \({\mathcal{O}}({\delta }^{2})\) good predictor, but at the same implementation cost as a single sub-unitary QNN.

In the rest of this work, we do not primarily take the QNN → LCU QNN → density QNN route as implied by Lemma 1, instead directly train the parameters of the density QNN indicating the viability of bypassing the expensive LCU directly. Though, for the sake of completeness, in Supplementary Material C we do provide one example where the LCU QNN outperforms the single-circuit QNN numercially. It may be that non-unitary (pure LCU) quantum machine learning35 and the density QNNs we present here have different difficulties in practice to achieve good model performance. Also, the above result is clearly restrictive to the setting where the model to be learned is itself a QNN, and we also know the correct observable to measure. However, it may be generalised to stronger statements where the data is generated by an arbitrary classical function. We leave both of these investigations to future work.

In the quantum compiling problem, one endeavours to produce a sequence of ‘simple’ operations which approximate the behaviour of a target unitary or quantum channel on input states. Randomised compiling (informally via the Hastings-Campbell mixing lemma40,41) allows one to carry a quadratic suppression in compiling error from a channel composed of a LCU, into the corresponding randomised channel which has the same execution overhead as an individual unitary in the linear combination. If we treat the compilation task as a learning problem, as has been done in several works, spawning the subfield of variational quantum compiling42,43, one can intuitively see how the Mixing lemma may be directly applied.

However, to argue for the benefit of density networks over individual QML circuits, we must go further and generalise to other learning tasks, for example supervised data classification/regression (which is the primary task we use density QNNs for in this work).

Connection to classical mechanisms

We have discussed features of density QNNs, and elucidated their position within the spectrum of existing quantum models. In this section, we discuss relationships or analogies between density QNNs to purely classical mechanisms and models. We summarise these relationships in the following three observations:

  • Observation 1: The randomised state preparation method may be used as the training mode, and the deterministic state preparation method may be used as the inference mode for density QNNs, if analogies are to be made with the classical dropout mechanism.

  • Observation 2: Unlike classical dropout, density QNNs do not combine an exponential number of sub-networks at inference time.

  • Observation 3: Density QNNs are a quantum analogue of Mixture of Experts (MoE) models, which are subtly different from ensemble methods.

We expand on these observations in the following.

Quantum dropout

In the previous section, we have demonstrated how density QNNs have the capacity to mitigate overfitting via more strategic parameter allocation. In light of this, one may make an analogy with the randomised density QNN and the dropout44,45 mechanism in classical neural networks. Dropout is an effective method of combining the predictions of exponentially many classical neural networks, and also mitigates overfitting. This analogy has indeed been remarked in recent works46 due to the random ‘removal’ of K − 1 sub-unitaries in each forward pass, which, on the surface, appears similar to the randomised removal of neurons in a neural network. it has therefore been conjectured that randomised density QNNs as we have defined them may be less prone to overfitting because of this analogy. However, in the following, we will argue that this is an incorrect, or at least an incomplete comparison.

This incompleteness of the comparison arises from (at least) two sources. Firstly, there is a distinct difference (and important) between the training and evaluation modes in classical dropout. In the training mode, a dropout mechanism applies a random zeroing of the nodes of a neural network, which removes outgoing weights from those nodes. This means, on any forward pass through the loss function, only a sub-network is actually activated. However, in the inference mode, all sub-networks are effectively present, where the outgoing weights of a dropped node are re-weighted by the probability of dropping that node.

Regarding different training and inference operations, we propose the randomised implementation of the density QNN (Fig. 1c) as the state preparation method for the training phase of the model. To align with classical dropout, we then propose the deterministic density QNN (Fig. 1b) as the inference mode. Due to the linearity of the model, this has the effect of weighting the contributions of each sub-unitary QNN, Uk(θk) by the corresponding coefficients, αk, exactly as in classical dropout.

However, dropout also has the key feature of efficiently combining an exponential number of effective sub-networks at inference, which is a critical feature in boosting model performance. However, in order to maintain training efficiency (within Corollary 2), a density QNN may have at most \(K={\mathcal{O}}(\log (N))\) ‘sub-networks’, where N is the total number of circuit parameters. As such, it is unclear if at a fundamental level a density QNN can function fully as a quantum analogue for dropout. Nevertheless, we will demonstrate that the density QNN still can, in some capacity, mitigate overfitting, and perform the effective action of dropout.

Density quantum neural networks as a mixture of experts

The next, and final, comparison to classical methods we discuss is an interpretation of the density QNN framework as a MoE model47,48. MoE models have achieved success even at the level of large scale machine learning models such as Mixtral49, particularly where sparsity is required. The MoE framework is flexible and powerful as it allows an increase in the capacity of a model with minimal increase in computational effort. This same methodology drives the density QNN framework of this work. A MoE contains a set of ‘experts’, {f1, …, fK}, each of which is ‘responsible’ for a different training case. A gating network, \({\mathcal{G}}\), decides which expert should be used for a given input. In the simplest form, the MoE output, F(x), is a weighted sum of the experts, according to the gating network output \(F({\boldsymbol{x}})=\sum _{k}{{\mathcal{G}}}_{k}({\boldsymbol{x}}){f}_{k}({\boldsymbol{x}})\) where \({{\mathcal{G}}}_{k}({\boldsymbol{x}})\) is the kth output of the gating network, given input x. A typical difference between MoE models and ensemble models is that, for the latter every element of the ensemble is evaluated for every input, while for an MoE, only a subset are activated (e.g. the top-k experts for a given input). We elaborate on this in Supplementary Material H.

We can interpret the density QNN exactly as a form of MoE as follows. By allowing the sub-unitary coefficients to be data-dependent (αkαk(x)), as in equation (3), we can predict them using a gating network in an MoE. Now, the sub-unitaries become the ‘experts’ and the overall model has increased capacity to learn which sub-unitary (expert) is more relevant for the task at hand (the given input x). If the data were quantum states, we could use another quantum neural network as a gating network, or perhaps classical shadows32 with a classical neural network. We choose a simple yet traditional model for the data-dependent gating mechanism we describe in ‘Methods’. There has been extensive investigation into variations of the MoE models both for deep50,51 and shallow models52,53 though we leave thorough investigation of different approaches in the quantum world to future work. For example, one could employ techniques regulating the dependence on the overall mixture on any individual subsets of unitaries50 which can occur when the distribution tends to become quite peaked, meaning the gating network chooses to rely only on a small fraction of experts for a given input.

Numerical results

Now, we demonstrate that density networks can be successful in practice, through three examples. The first is an demonstration where an efficiently trainable model can be made more expressive (see Fig. 3, right), and in the second we choose an example of an interpretable model family for which we construct both efficiently trainable versions (Fig. 3, left) and more expressive (Fig. 3, right). The third, and final, main example demonstrates the ability of density QNNs to mitigate overfitting compared to standard single unitary QML models. We summarise the takeaways of each set of numerics as follows:

  • Model: Equivariant quantum neural networks.

  • Model: Hamming-weight preserving (orthogonal) quantum neural networks.

  • Model: Data reuploading quantum neural networks.

Equivariant quantum neural networks

For the first example, we build a density version of a commuting-generator (a special case of commuting-block circuit) ansatz on the simplified classification problem posed by ref. 38. Here, the challenge is classifying bars vs. dots, in a noisy setting, described in ‘Methods’.

For the base QNN, we use the original ‘XX’ ansatz (denoted UXX(θ)) of ref. 38, see Fig. 4a) and compare against a density QNN version with two sub-unitaries, {U1, U2} and coefficients {α1, α2}. We define U1 UXX(θ). The second sub-unitary, U2 UYY(θ), has the exact same structure as UXX(θ), but with the XX generators replaced by Pauli-Y generators. The measurement operator is chosen to be the translation invariant operator \(\sum _{i}{Z}_{i}\) to suit the translation symmetry in the synthetic problem. The output function of the model is then:

$$f({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}}):= \mathop{\sum }\limits_{i=1}^{d}\mathop{\sum }\limits_{k=1}^{K}{\rm{Tr}}({Z}_{i}\rho ({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}}))$$

(11)

with

$$\begin{array}{rcl}\rho ({\boldsymbol{\theta }},{\boldsymbol{\alpha }},{\boldsymbol{x}})&:= &{\alpha }_{1}{U}_{XX}({\boldsymbol{\theta }})\left[\mathop{\bigotimes }\limits_{j=1}^{d}\left\vert {\boldsymbol{x}}\right\rangle {\left\langle {\boldsymbol{x}}\right\vert }_{j}^{ry}\right]{U}_{XX}^{\dagger }({\boldsymbol{\theta }})\\ &&+{\alpha }_{2}{U}_{YY}({\boldsymbol{\theta }})\left[\mathop{\bigotimes }\limits_{j=1}^{d}\left\vert {\boldsymbol{x}}\right\rangle {\left\langle {\boldsymbol{x}}\right\vert }_{j}^{ry}\right]{U}_{YY}^{\dagger }({\boldsymbol{\theta }})\end{array}$$

(12)

Fig. 4: Equivariant density QNNs.
figure 4

a The commuting-generator XX model and a, b XX+YY density QNN model. The former contains up to three-body Pauli-X generated operations with twirling applied to enforce equivariance. The latter contains two sub-unitaries UXX (circuit (a)) and UYY which has the same structure but replacing Pauli-X operations with Pauli-Y. UXX/UYY are applied with probabilities αXX/YY. Each sub-unitary in b are commuting-generator circuits, so each has efficiently extractable gradients.

Each generator that appears in UXX/YY(θ) is of the form \({\mathsf{sym}}({X}_{1})\) or \({\mathsf{sym}}({X}_{1}\ldots {X}_{k})\) (\({\mathsf{sym}}({Y}_{1})\) and \({\mathsf{sym}}({Y}_{1}\ldots {Y}_{k})\) respectively) for some k ≤ K. The operation \({\mathsf{sym}}\) is the twirling operation used to generate equivariant quantum circuits (in this case, equivariance with respect to translation symmetry). For example, \({\mathsf{sym}}({X}_{1}{X}_{2})\) is a sum of all pairs of X operators on the state with no intermediate trivial qubit, \({\mathsf{sym}}({X}_{1}{X}_{3})\) is a sum of all pairs on the state separated by exactly one trivial qubit (in either direction, visualising the qubits as a 1D chain with closed boundary conditions) and so on. However, note that in this case since the YY ansatz (UXX/YY(θ)) is in the same basis as the data encoding (Ry), we have classical simulatability for computing expectation values since the initial circuit is effectively Clifford38. The density circuits are visualised in Fig. 4b which we adapt from ref. 38, and the results of the experiment can be seen in Fig. 5a, b. We also construct density versions of the two other non trivial models considered in ref. 38, the quantum convolutional neural network (QCNN)54 and a ‘non-commuting` equivariant circuit (see ref. 38, Fig. 5c). In both these latter cases, we again consider two sub-unitaries U1 and U2 which have identical structure but different learned parameterisations, U1 = U2, \({{\boldsymbol{\theta }}}_{1}^{* }\ne {{\boldsymbol{\theta }}}_{2}^{* }\). In initialising the density QNN, we bias the model towards the ‘first’ sub-unitary with probability 99% (as in Fig. 4b). The density QNN outperforms its base counterpart in all cases, and the best tradeoff between efficiency and model performance is found for the density XX + YY model. We describe the experiment details in ‘Methods’.

Fig. 5: Numerics on noisy bars and dots dataset.
figure 5

We create density QNN versions of the non-trivial models from ref. 38; (1) the commuting-generator circuit in Fig. 4, (2) a `non-commuting’ equivariant QNN and (3) the quantum convolutional neural network54, all on 10 qubits. In all cases the density QNN is initialised from (separately) pretrained (for 1000 epochs) base versions, and training continues for another 1000. For the non-commuting and QCNN density models the sub-unitaries have identical structure but trained independently. We show mean and standard deviation in test accuracy vs. a training epochs and b number of overall shots over 5 independent training runs, from the same initialisation. In all cases, after base model performance saturation, the density QNN improves the final result. The gaps in (b) are to account for the extra measurement overhead to initialise and train the second sub-unitaries. This is UYY for the density model and U2 (non-comm-2/QCNN-2) for the other two models. In all cases, the density version is initialised with α1/αXX = 0.99 α2/αYY = 0.01 which are also trainable.

Orthogonal quantum neural networks

For the second example, we turn to the Hamming weight (HW) preserving quantum neural network (U(1) equivariant), and specifically their orthogonal QNNs variants. As mentioned above, these models have desirable properties from a machine learning point of view—they are interpretable and can stabilise training. While we focus on the easiest to simulate version of HW preserving models here, all of the below is applicable to the more difficult to classically simulate compound QNNs55 (Supplementary Material E.2) and general Hamming weight preserving unitaries.

The data encoded state for orthogonal QNNs on n qubits, \(\rho ({\boldsymbol{x}}):= \left\vert {\boldsymbol{x}}\right\rangle \left\langle {\boldsymbol{x}}\right\vert ,\left\vert {\boldsymbol{x}}\right\rangle := \sum _{j}{x}_{j}\left\vert {{\bf{e}}}_{j}\right\rangle\), is a unary amplitude encoding of the vector x (ej is a basis vector with a single 1 in position j and zeros otherwise). The unitary that acts on this state consists of two qubit gates known as reconfigurable beam splitter (RBS) gates (see ‘Methods’). There are various configurations for these gates which parameterise some subset of SO(n) matrices. To begin, we choose the round-robin architecture, shown in Fig. 6 (top), composed of \({\mathcal{O}}(n)\) layers of where in each layer qubit is connected to only one other, alternating so eventually each qubit is connected to every other. Evaluating gradients for this QNN specification via the parameter shift rule requires \({\mathcal{O}}({n}^{2})\) circuits. We construct the ‘shallow’ version of a round-robin density QNN by simply extracting each layer into the sub-unitaries, with \(K={\mathcal{O}}(n)\). Then each Uk is a depth D = 1 commuting generator circuit whose gradients can be evaluated in parallel with \({\mathcal{O}}(n)\) circuits via Corollary 2. In the spirit of Fig. 3 (right), we also take \(K={\mathcal{O}}(n)\) copies of the full round robin circuit to construct a more expressive model. This model will require \({\mathcal{O}}(K{n}^{2})={\mathcal{O}}({n}^{3})\) circuits for gradient evaluation, but also has \({\mathcal{O}}(n)\) more parameters than the ‘vanilla’ round-robin model. We can also interpolate between these two regimes by scaling the complexity of each sub-unitary with a round-robin layer depth between 1 ≤ D ≤ n − 1. In all cases, we demonstrate the MoE formalism by introducing a gating network to predict the coefficients, {α(x)}, details given in ‘Methods’.

Fig. 6: Hamming weight preserving density QNN with round-robin structure.
figure 6

Illustration of extraction possibilities from an orthogonal QNN with round-robin connectivity, on 8 qubits. We can extract a minimally expressive `shallow’ (depth D = 1) or a maximally expressive `deep’ (depth \(D=\frac{n(n-1)}{2}\)) sub-unitary for each ‘expert’. Both cases have K = n − 1 sub-unitaries, which have \(\frac{n}{2}\) (shallow) or \(\frac{n(n-1)}{2}\) (deep) parameters. The shallow version has gradients which can be evaluated in parallel using n − 1 circuits with gradient complexity stated in Table 1. The deep version assumes a cycled connectivity as seen in the figure with gradient extraction complexity \({\mathcal{O}}({n}^{3})\). We can interpolate between the two regimes for different sub-unitary depths, D, between 1 ≤ D ≤ n − 1, but efficient gradient extraction is lost for D ≥ 2. An MoE gating network can be used to predict the probability coefficients, {αk(x)}.

In all cases, since the unitaries are Hamming weight preserving, the output states, \(\left\vert {{\boldsymbol{y}}}^{k}\right\rangle\) from each sub-unitary are of the form \(\vert {\boldsymbol{y}}\rangle =\sum _{j}{y}_{j}\vert {{\bf{e}}}_{j}\rangle\) for some vector y. This output is related to the input vector via some orthogonal matrix transformation OU, y = OUx where the elements of OU can be computed via the angles of the RBS gates in the circuit (see ‘Methods’). The typical output of such a layer is the vector y itself, for further processing in a deep learning pipeline. As a result the output of a ‘orthogonal’ density QNN is a linear combination of orthogonal transformations, \({\boldsymbol{y}}=\sum _{k}{\alpha }_{k}{{\boldsymbol{y}}}_{k}=\sum _{k}{\alpha }_{k}{O}^{{U}_{k}}{\boldsymbol{x}}\), benefitting from both stable orthogonal training within each sub-unitary, and the generality afforded by the linear combination. In contrast, adding more (e.g. RBS) gates directly to the pure state version circuit (i.e. the ‘vanilla’ round-robin QNN) would not increase the expressivity as the output would only correspond to another (different) orthogonal matrix. The results of these experiments are shown in Fig. 7.

Fig. 7: Round-robin Hamming weight preserving density QNN results.
figure 7

Top row Round-robin OrthoQNN versus shallow (D = 1) round-robin density QNN, as a function of training epochs (a) and measurement shots (b) on 10 qubits. For D = 1 the density QNN is a commuting generator circuit and so requires only n − 1 circuit evaluations for each iteration whereas the OrthoQNN requires \(\frac{4n(n-1)}{2}\) circuits as per the parameter-shift rule for gradient evaluation. Bottom row increasing complexity of density QNN by increasing round robin depth, D, from D {1, 2, 3, 4, (n − 1)} as a function of epochs (c) and measurement shots (b, log scale). In all cases for the density QNN, n − 1 = 9 sub-unitaries are trained in parallel for 25 epochs (9 thin lines) before initialising the density QNN, and training continues for a further 25. Performance of density QNN increases monotonically with sub-unitary depth, D.

Mitigating overfitting with density QNNs

We have argued above that the density QNN framework possesses similarities and differences to the dropout mechanism. Nevertheless, we can demonstrate that indeed, density QNNs have the capacity to do what dropout does—which is to mitigate overfitting. To demonstrate this, we use data reuploading56 QNNs. Data reuploading is a poweful QML technique which allows the construction of universal quantum classifiers, even with single qubits. We compare a single, deep (‘vanilla’) data reuploading QNN to a density version with approximately the same number of parameters, and test which model is more likely to overfit training data, generated by Chebyshev polynomials—see ‘Methods’ for details. In summary, we find that given a budget of N parameters, distributing these across a density QNN with shallow sub-unitaries will lead to better results (from the perspective of overfitting) than increasing the circuit depth of a single quantum classifier. See ‘Methods’ for a definition of data reuploading and Supplementary Material G.3 for a discussion of data reuploading with density QNNs.

For both density and vanilla reuploading models, we use arbitrary single qubit rotations with 3 parameters as the trainable operations. For the vanilla version, these are repeated for L ‘reuploads’ so the model has 3L parameters. For the density QNN, we choose K = 5 sub-unitaries, the depth of each (D) is chosen such that 5 × 3D ≈ 3L. Both models are illustrated in Fig. 8a, and we also shown the truncated Fourier series learned by each of the 5 sub-unitaries in the density case. The results can be seen in Fig. 8b which shows the test error, as measured by mean squared error (MSE) and the generalisation gap (difference between train and test errors).

Fig. 8: Overfitting mitigation with density QNNs—data reuploading.
figure 8

a A ‘vanilla’ single qubit data reuploading QNN with L reuploading layers versus a density QNN with 5 sub-unitaries, each with D data reuploads. Also shown is the partial Fourier series representations of each sub-unitary, \({\{\,{f}^{k}({{\boldsymbol{\theta }}}_{k},x)\}}_{k = 1}^{5}\) when learning the Chebyshev polynomial T2(x). b Generalisation gap between train and test mean square errors (MSE) in a regression problem for learning the Chebyshev polynomial T2(x). The number of parameters between both models is kept approximately the same for each L, so that 5 × 3D ≈ 3L. For a fixed number of parameters, the test error and train/test gap (generalisation error) for the vanilla data reuploading model diverges, while the density QNN test error remains small.

We see as the vanilla reuploading model becomes more expressive (with increasing L), it overfits and test error grows. In contrast, the density model (with approximately the same number of parameters) does not overfit and the test error remains small. This means that if, without any prior, one would choose an overparameterised model (see57,58 for further discussions on overparameterisation which may be incorporated in future work), the density model would perform better. In Supplementary Material J we provide further detail on this, and supplementary numerics.

Drawbacks and limitations

To conclude our results, we discuss some potential drawbacks and limitations with the density framework, specifically those which it does not solve relative to single-unitary variational learning models. A major problem with the latter is the existence of barren plateaus, or regions of problematic gradients more generally. As discussed in ref. 38 the general relationship between commuting generator circuits, and barren plateaus is still an open question—particularly in finding a circuit architecture which has a suitable relationship between input states, circuit generators and measurement observable to render the dynamical Lie algebra polynomial, the circuit non-classically simulatable and gradients non-vanishing. This is not a question resolved via the density framework either, although if such an architecture was to be found, it could similarly be uplifted as other specific models studied in this work. However, an interesting research direction provided by the density framework is to allow a combination of different sub-unitaries, which may have different and independent gradient behaviour. However, in this case it is possible that the sub-unitary with most favourable gradients would dominate the training (a feature known as representation or expert collapse in the mixture of expert literature59,60,61). However, it has also been demonstrated how barren plateau issues can be mitigated or eliminated via clever initialisation strategies.

Besides the barren plateau issue, there are still a number of open questions with variational quantum learning models, which are also potentially inherited by non-unitary (LCU, dissipative or density) learning models including; the effect of quantum or measurement noise, lack of data or problem-dependent circuit architectures and training dynamics related to hyperparameter tuning and loss function choices. All of these are fruitful areas for future study.



Source link