Telling Two Quantum Channels Apart
How well can you tell two quantum channels apart? Suppose a black box applies either channel $\Phi_0$ or channel $\Phi_1$, each with probability $1/2$. You may feed it any input, including half of an entangled state, and measure the output however you like. The best possible success probability is
$$
P_{\text{succ}}^{\max}=\frac12+\frac14,\big|\Phi_0-\Phi_1\big|_{\diamond},
$$
where $|\cdot|_\diamond$ is the diamond norm. This post defines the norm, turns it into a semidefinite program (SDP), and solves it in Python for concrete channels.
1. The diamond norm
For a Hermiticity-preserving map $\Delta:\mathcal L(\mathcal H_{\text{in}})\to\mathcal L(\mathcal H_{\text{out}})$, the diamond norm is

where the maximum runs over all states on an input system $A$ plus a reference (ancilla) system $R$. The trace norm is $|X|_1=\mathrm{Tr}\sqrt{X^\dagger X}$.
The ancilla matters. A channel may look identical to another on every single input state and still be distinguishable once the input is entangled with a reference.
2. The Choi matrix
For a channel $\Phi$ with $d_{\text{in}}=d_{\text{out}}=d$, define
$$
J(\Phi)=\sum_{i,j=1}^{d}|i\rangle\langle j|\otimes\Phi\big(|i\rangle\langle j|\big).
$$
Then $J(\Phi)$ is positive semidefinite iff $\Phi$ is completely positive. For a difference of two channels, $J(\Delta)=J(\Phi_0)-J(\Phi_1)$ is a Hermitian matrix. It is generally indefinite.
3. The SDP (Watrous’ formulation)
The diamond norm of $\Delta$ is the optimal value of

The variables $Y_0,Y_1$ are Hermitian matrices of size $d_{\text{in}}d_{\text{out}}$, and $t_0,t_1$ are real scalars. The two scalar constraints are exactly the statements $|\mathrm{Tr}_{\text{out}}Y_k|_\infty\le t_k$, because the $Y_k$ are positive semidefinite. Everything is linear in the unknowns, so this is a standard SDP.
4. The example problems
Example 1: identity versus depolarizing. The depolarizing channel is $\mathcal D_p(\rho)=(1-p)\rho+p,\mathrm{Tr}(\rho),\mathbb 1/2$. Then $\mathrm{id}-\mathcal D_p=p,(\mathrm{id}-\mathcal D_1)$. The optimal input is the maximally entangled state $|\Phi^+\rangle$, and
$$
|\mathrm{id}-\mathcal D_p|_\diamond=p,\Big|,|\Phi^+\rangle\langle\Phi^+|-\tfrac{\mathbb 1}{4}\Big|_1=\frac{3p}{2}.
$$
Example 2: identity versus a $Z$-rotation. Let $\mathcal R_\theta(\rho)=U_\theta\rho U_\theta^\dagger$ with $U_\theta=\mathrm{diag}(e^{-i\theta/2},e^{i\theta/2})$. The difference of two unitary channels has the closed form

(the minimum is over density matrices). For the $Z$-rotation this evaluates to
$$
|\mathrm{id}-\mathcal R_\theta|_\diamond=2,\big|\sin(\theta/2)\big| .
$$
Example 3: depolarizing versus amplitude damping. The amplitude damping channel has Kraus operators
$$
K_0=\begin{pmatrix}1&0\0&\sqrt{1-\gamma}\end{pmatrix},\qquad
K_1=\begin{pmatrix}0&\sqrt{\gamma}\0&0\end{pmatrix}.
$$
No simple formula exists for $|\mathcal D_p-\mathcal A_\gamma|_\diamond$. We compute it on a grid of $(p,\gamma)$ and read off the optimal success probability $\tfrac12+\tfrac14|\mathcal D_p-\mathcal A_\gamma|_\diamond$.
Entanglement advantage. For comparison we also compute the best distinguishability achievable without an ancilla,
$$
\max_{|\psi\rangle}\big|\Phi_0(\psi)-\Phi_1(\psi)\big|_1,
$$
by sampling the Bloch sphere. The gap to the diamond norm is the benefit of entangled inputs.
5. Source code
1 | import time |
6. Code walkthrough
6.1 Channels and Choi matrices
Each channel is a Python function acting on a $2\times2$ matrix. Because every channel is linear, the same function works on density matrices and on the matrix units $|i\rangle\langle j|$ that appear in the Choi construction. choi() builds
$$
J(\Phi)=\sum_{i,j}|i\rangle\langle j|\otimes\Phi(|i\rangle\langle j|)
$$
with np.kron(E, channel(E)). The input system is the first tensor factor and the output system is the second. The SDP uses the same ordering, which is why a partial trace over the second factor is the right operation.
6.2 The SDP in diamond_norm
Y0andY1are Hermitian CVXPY variables of size $4\times4$ for a qubit.blockis the $8\times8$ matrix $\begin{pmatrix}Y_0&-J\-J^\dagger&Y_1\end{pmatrix}$, andblock >> 0is the positive-semidefinite constraint.- The partial trace over the output is written without any special atom. For each output basis vector $|k\rangle$ we form the constant matrix $A_k=\mathbb 1_{\text{in}}\otimes\langle k|$, and then
$$
\mathrm{Tr}_{\text{out}},Y=\sum_k A_k,Y,A_k^{\mathsf T}.
$$
This keeps the model purely affine in the variables and works for complex Hermitian matrices. ptrace_out(Y0) << t0*Iencodes $\mathrm{Tr}_{\text{out}}Y_0\preceq t_0\mathbb 1$, the standard way to express a spectral-norm bound inside an SDP.- The objective is $\tfrac12(t_0+t_1)$. The solver is Clarabel, an interior-point method that reaches high accuracy on small SDPs, with SCS as a fallback.
6.3 Speed
Each SDP involves only $4\times4$ and $8\times8$ matrices. The optimal values are obtained to roughly $10^{-8}$ accuracy by an interior-point solver in a few milliseconds each, so the whole study of about 280 SDPs finishes in seconds. The Bloch-sphere scan uses only $2\times2$ eigenvalue computations. The 3D map therefore needs no approximation scheme, only a modest $13\times13$ grid that is enough for a smooth surface.
6.4 Comparison tools
trace_norm() computes $|M|_1$ from the eigenvalues of the Hermitian matrix $M$. best_single_input() scans pure input states $|\psi(\theta,\phi)\rangle$ over the Bloch sphere and returns $\max_\psi|\Phi_0(\psi)-\Phi_1(\psi)|_1$, the best any ancilla-free strategy can do.
7. Results
Execution result image

Console output
Solver: CLARABEL
================================================================
(A) identity vs depolarizing(p) [analytic: 3p/2]
p SDP analytic abs.err
0.00 0.000000 0.000000 3.77e-11
0.10 0.150000 0.150000 2.78e-10
0.20 0.300000 0.300000 1.75e-10
0.30 0.450000 0.450000 1.79e-10
0.40 0.600000 0.600000 1.14e-10
0.50 0.750000 0.750000 2.19e-09
0.60 0.900000 0.900000 5.60e-09
0.70 1.050000 1.050000 1.17e-08
0.80 1.200000 1.200000 8.65e-09
0.90 1.350000 1.350000 3.62e-09
1.00 1.500000 1.500000 2.08e-09
max error = 1.17e-08
================================================================
(B) identity vs Z-rotation(theta) [analytic: 2|sin(theta/2)|]
theta/pi SDP analytic abs.err
0.000 0.000000 0.000000 3.77e-11
0.333 1.000000 1.000000 6.22e-10
0.667 1.732051 1.732051 5.52e-10
1.000 2.000000 2.000000 2.20e-09
1.333 1.732051 1.732051 5.52e-10
1.667 1.000000 1.000000 6.22e-10
2.000 0.000000 0.000000 6.39e-12
max error = 2.20e-09
================================================================
(C) depolarizing vs amplitude damping on a 13x13 grid
max diamond distance = 2.000000 at p = 0.000, gamma = 1.000
Helstrom success probability range: 0.5000 .. 1.0000
================================================================
(D) fixed pair p = 0.6, gamma = 0.3
diamond norm (SDP, entangled input) = 0.665150
best single-input trace norm (sphere) = 0.600936
================================================================
(E) single input vs ancilla-assisted (diamond norm)
id vs full depolarizing single = 1.000000 diamond = 1.500000
id vs Z-dephasing single = 2.000000 diamond = 2.000000
depol(0.6) vs ampl.damp(0.3) single = 0.600941 diamond = 0.665150
================================================================
Total elapsed time: 20.2 s
8. Reading the figure
Panel (A): identity versus depolarizing. The SDP values lie on the line $3p/2$. At $p=1$ the value $3/2$ is the trace distance between the maximally entangled state and the maximally mixed state on two qubits, $|,|\Phi^+\rangle\langle\Phi^+|-\mathbb 1/4|_1=\tfrac34+3\cdot\tfrac14$. The match with the analytic line shows that the SDP, the Choi convention and the partial trace are all correct.
Panel (B): identity versus $Z$-rotation. The SDP reproduces $2|\sin(\theta/2)|$. The curve peaks at $\theta=\pi$, where $U_\theta^\dagger$ has eigenvalues $\pm i$, so the convex hull of the eigenvalues contains the origin. There the two unitaries are perfectly distinguishable and the norm reaches its maximum value $2$. At $\theta=2\pi$ the unitary equals $-\mathbb 1$, which is the identity channel up to a global phase, so the norm returns to $0$.
Panel (C): the 3D landscape. For depolarizing versus amplitude damping, the surface is zero only where the two channels coincide, namely the corner $p=\gamma=0$ where both are the identity. It rises toward the corner $p=0,\ \gamma=1$, where the norm reaches its maximum of $2$. There the identity channel is compared with the channel that resets every state to $|0\rangle$, and these two outputs can be made orthogonal. The surface is a continuous, convex-looking sheet, as expected, since the diamond norm is a norm and both channels depend affinely on $p$ and $\gamma$ (the Kraus-form amplitude damping channel is affine in $\gamma$ at the level of its action on density matrices).
Panel (D): single-input distinguishability on the Bloch sphere. Each point of the sphere is a pure input state, coloured by $|\Phi_0(\psi)-\Phi_1(\psi)|_1$ for $p=0.6,\ \gamma=0.3$. The maximum over the sphere is smaller than the diamond norm printed in the title. The remaining difference can be obtained only by sending half of an entangled state and measuring the pair jointly.
Panel (E): optimal success probability. The contour map shows $\tfrac12+\tfrac14|\Phi_0-\Phi_1|_\diamond$ over the $(p,\gamma)$ plane. Success ranges from $1/2$, which is a coin flip, at the coincident corner up to $1$ at the corner where the channels are perfectly distinguishable. The star marks the pair studied in panels (D) and (F).
Panel (F): the entanglement advantage. Three cases appear side by side.
- Identity versus full depolarizing: the best ancilla-free value is $1$, but the diamond norm is $3/2$. This is the clean example of a gap created by entanglement.
- Identity versus $Z$-dephasing with $p=1$: both values equal $2$. A single input such as $|+\rangle$ is already perfectly sufficient, since $|+\rangle$ and $|-\rangle$ are orthogonal, so entanglement brings no improvement.
- Depolarizing versus amplitude damping at $(0.6,0.3)$: the diamond norm exceeds the best single-input value, a smaller but genuine advantage.
9. Takeaways
- The diamond norm is the correct figure of merit for channel discrimination, because it accounts for entangled inputs and adaptive measurements.
- A semidefinite program with only two small positive-semidefinite blocks computes it exactly for qubit channels.
- Closed-form cases (depolarizing noise and unitary rotations) agree with the SDP to numerical precision, which makes them good unit tests for any implementation.
- Optimal discrimination can require an ancilla. The gap between the best single-input value and the diamond norm is the quantitative benefit of entanglement.
The same code extends to larger systems by changing the dimension arguments d_in and d_out and supplying the Choi matrix of the difference of two channels.





. The dark-to-bright scale makes solver-level errors visible.





first and are then pushed further by the logarithmic growth of $x^{*}$. Expensive resources appear mainly at the ridge peaks, and cheap resources also cover the quiet hours.

