Lanthaler_2026_TopologicalDeepONets_arXiv_2603.11972

📄 来源 PDF:Lanthaler_2026_TopologicalDeepONets_arXiv_2603.11972.pdf

正文(按页提取)

第 1 页

6
2
0
2

r
a

M
2
1

]

G
L
.
s
c
[

1
v
2
7
9
1
1
.
3
0
6
2
:
v
i
X
r
a

Topological DeepONets and a generalization
of the Chen–Chen operator approximation
theorem

Vugar E. Ismailov∗

Abstract

Deep Operator Networks (DeepONets) provide a branch–trunk neural architecture
for approximating nonlinear operators acting between function spaces. In the clas-
sical operator approximation framework, the input is a function u ∈ C(K1) defined
on a compact set K1 (typically a compact subset of a Banach space), and the oper-
ator maps u to an output function G(u) ∈ C(K2) defined on a compact Euclidean
domain K2 ⊂ Rd. In this paper, we develop a topological extension in which the
operator input lies in an arbitrary Hausdorff locally convex space X. We construct
topological feedforward neural networks on X using continuous linear functionals
from the dual space X ∗ and introduce topological DeepONets whose branch com-
ponent acts on X through such linear measurements, while the trunk component
acts on the Euclidean output domain. Our main theorem shows that continuous
operators G : V → C(K; Rm), where V ⊂ X and K ⊂ Rd are compact, can be
uniformly approximated by such topological DeepONets. This extends the classical
Chen–Chen operator approximation theorem from spaces of continuous functions
to locally convex spaces and yields a branch–trunk approximation theorem beyond
the Banach-space setting.

Keywords: DeepONet; topological neural network; locally convex space; branch–

trunk network; universal approximation theorem; operator approximation.

2020 MSC: 68T07, 41A30, 41A65, 46A03, 47H99

1

Introduction

Deep neural networks are usually used to approximate nonlinear mappings between finite-
dimensional Euclidean spaces. However, in many scientific and engineering applications,
the object of interest is not a function but an operator: a mapping that takes an input
function and returns an output function. In the operator-learning viewpoint, one aims
to learn a continuous nonlinear operator

G : V −→ C(K; Rm),

from sampled input–output data, where V is a compact set of admissible inputs and K
is the output domain.

∗The author can be contacted at vugaris@mail.ru or vugaris@gmail.com.

1

第 2 页

A prominent architecture for this purpose is the Deep Operator Network (DeepONet),
proposed in [14]. In its standard form, DeepONet employs a branch network (encoding
the input function through sensor measurements) and a trunk network (processing the
variable y), and combines their outputs through a dot product. In this way, DeepONet
approximates G(u)(y) in a separable form as a finite sum of functions of y, with coefficients
depending on the input u.

A central theoretical motivation for DeepONets comes from the universal approxima-
tion theorem for operators established by Chen and Chen [3]. This theorem shows that
continuous operators between spaces of continuous functions can be uniformly approxi-
mated on compact sets by expressions depending on finitely many point evaluations of
the input function. DeepONets place such approximants into a structured branch–trunk
architecture and, in practice, make both sides deep.

Since its introduction, DeepONet has been successfully applied to a wide range of engi-
neering and scientific problems. Early studies demonstrated its ability to learn operators
associated with dynamical systems and diffusion–reaction equations [3, 14]. Subsequent
works employed DeepONets for learning nonlinear operators arising in multiphysics and
multiscale models, including electro-convection phenomena [2], hypersonic flows governed
by the Navier–Stokes equations with finite-rate chemistry [15], and multiscale bubble-
growth dynamics [13]. Related developments include physics-informed DeepONets for
learning solution operators of parametric PDEs [23].

More recent developments extend DeepONet-based operator learning to a variety
including complex-valued formulations for three-
of challenging application domains,
dimensional Maxwell equations [8], surrogate modeling for shape optimization [20], learn-
ing two-phase microstructure evolution using neural-operator and autoencoder architec-
tures [16], and aerothermodynamic analysis of hypersonic configurations [21]. In addi-
tion, DeepONet architectures have been explored in control-oriented settings, including
neural-operator approximation of backstepping controller and observer gain functions
for reaction–diffusion PDEs [10], predictor-based stabilization of nonlinear systems with
input delay [1], and predictive control [9].

The aim of the present paper is to develop a framework in which the input of the
operator need not lie in a Euclidean space or, more generally, in a normed linear space
(such as a Hilbert or Banach space). Instead, we treat the case where the input belongs
to a locally convex topological vector space X. In this viewpoint, the network receives
admissible measurements of the input element in the form of continuous linear functionals.
That is, each hidden neuron evaluates a continuous linear functional on the input element
and then applies the activation function. This setting is natural when the input is an
element of an abstract function space endowed with a locally convex topology and the
architecture is allowed to use linear measurements compatible with that topology. Such
situations arise frequently in analysis and applications.

For instance, spaces of differentiable functions, which arise naturally in the theory of
partial differential equations, provide fundamental examples of non-normable topological
spaces. The Schwartz space S(Rn) of rapidly decreasing smooth functions is equipped
with the countable family of seminorms

pa,b(f ) = sup
x∈Rn

(cid:12)xaDbf (x)(cid:12)
(cid:12)
(cid:12),

a, b ∈ Nn,

where xa = xa1
hence a Fr´echet space, but it is not normable.

n and Db = ∂|b|/∂xb1

1 · · · xan

1 · · · ∂xbn

n . This space is complete and metrizable,

2

第 3 页

Another important example is the space D(U ) of smooth functions with compact

support in an open set U ⊂ Rn. For each compact set K ⊂ U , the subspace

C ∞

0 (K) = {f ∈ C ∞(U ) : supp(f ) ⊂ K}

is a Fr´echet space with seminorms

pK,m(f ) = max
|α|≤m

sup
x∈K

|Dαf (x)|,

m ∈ N.

The space D(U ) is obtained as the inductive limit of the spaces C ∞
directed family of compact subsets of U whose union equals U , that is,

0 (K), where (Kj) is a

D(U ) = lim
−→
In this topology, D(U ) is locally convex and complete, but not metrizable and therefore
not normable. The space D(U ), known as the space of test functions, plays a fundamental
role in the theory of distributions.

0 (Kj).

C ∞

More generally, for a topological space X, the space C(X) of continuous functions
endowed with the topology of uniform convergence on compact sets, defined by the semi-
norms

φK(f ) = max
x∈K
is a locally convex space that is not normable unless X is compact.

K ⊂ X compact,

|f (x)|,

In our recent work [6] we proved a universal approximation theorem for feedforward
neural networks on topological vector spaces with the Hahn–Banach extension property,
in particular on locally convex spaces. These networks are constructed using continuous
linear functionals from the dual space X ∗ together with a fixed scalar activation function.
The present paper uses this density mechanism on compact sets to build a DeepONet-type
approximation framework and to prove an operator approximation theorem of branch–
trunk type.

Specifically, we introduce a topological DeepONet architecture in which the branch
component acts on a locally convex input space X through continuous linear mea-
surements from X ∗, while the trunk component acts on the output domain K ⊂ Rd.
Within this setting we prove a universal approximation theorem for continuous opera-
tors G : V → C(K; Rm) on compact sets V ⊂ X by finite separable expansions whose
coefficient maps are realized by topological neural networks on X.

As a consequence, we obtain approximation theorems related to DeepONets, showing
that continuous nonlinear operators admit approximations of branch–trunk type extend-
ing the classical dot-product formulation beyond the Banach-space setting. The results
place the Chen–Chen operator approximation principle [3] and the DeepONet architec-
ture [14] into a unified locally convex framework.

The remainder of the paper is organized as follows. Section 2 introduces topological
neural networks and recalls the universality theorem on locally convex spaces. Section 3
develops the topological DeepONet architecture and proves the main operator approxi-
mation theorems and their corollaries. Section 4 presents several examples.

2 Topological networks on locally convex spaces

This section provides definitions and results that will be used later.

Throughout the paper, X denotes a locally convex topological vector space and X ∗

its continuous dual. All locally convex spaces are assumed to be Hausdorff.

3

第 4 页

Definition 2.1 (Topological neural network on a locally convex space). Fix an activation
function σ : R → R and let m ≥ 1. A (vector-valued) topological feedforward neural
network on X with one hidden layer is any mapping H : X → Rm of the form

H(x) = A σ(cid:0)T (x)(cid:1),

where

T (x) = (f1(x) − θ1, . . . , fr(x) − θr),
with fi ∈ X ∗, θi ∈ R, and A ∈ Rm×r. Here σ acts componentwise on Rr. Each hidden
neuron evaluates a continuous linear functional on the input element and then applies
the activation function.

Deep networks are obtained by alternating affine maps and componentwise activation
functions, with the output layer given by a linear map. For completeness, we recall one
convenient formulation.

Definition 2.2 (Deep topological neural network on X). Fix integers L ≥ 2 and m ≥ 1,
and let n1, . . . , nL−1 ≥ 1. A (vector-valued) deep topological feedforward neural network
on X of depth L is a mapping H : X → Rm of the form

H(x) = AL zL−1(x),

where the hidden-layer outputs zℓ(x) ∈ Rnℓ are defined by
zℓ(x) = σ(cid:0)Aℓzℓ−1(x) − bℓ

z1(x) = σ(cid:0)T1(x)(cid:1),

(cid:1), ℓ = 2, . . . , L − 1.

Here σ acts componentwise, each Aℓ is a real matrix of appropriate size, and each bℓ is a
real bias vector. The first affine map T1 : X → Rn1 is defined by

T1(x) = (cid:0)f1(x) − θ1, . . . , fn1(x) − θn1

(cid:1),

with fj ∈ X ∗ and θj ∈ R.

Remark 2.1. When X = Rd endowed with its usual topology, the above classes of
networks reduce to the standard feedforward neural networks used in the classical theory.
Indeed, in this case every continuous linear functional f ∈ X ∗ has the form

f (x) = w · x,

x ∈ Rd,

for some vector w ∈ Rd, where w · x denotes the Euclidean inner product. This follows
from the well known representation theorem for continuous linear functionals on Hilbert
spaces. Consequently, the expressions fj(x) − θj appearing in the first affine map become

wj · x − θj,

which are exactly the affine forms used in classical neural networks on Rd. Therefore,
the topological neural networks introduced above extend the traditional Euclidean neural
networks to inputs belonging to general locally convex spaces.

4

第 5 页

We work with the space C(X; Rm) of continuous functions from X into Rm equipped
with the topology of uniform convergence on compact sets. This topology is generated
by the family of seminorms

∥g∥K = sup
x∈K
where K ranges over all compact subsets of X and ∥ · ∥Rm denotes any fixed norm on Rm.
Since all norms on a finite-dimensional space are equivalent, the resulting topology does
not depend on the particular choice of the norm.

∥g(x)∥Rm,

A subbasis at the origin for this topology is given by the sets

U (K, r) = {g ∈ C(X; Rm) : ∥g∥K < r} ,

where K ⊂ X is compact and r > 0.

Thus, when we say that a family of functions F acting from X into Rm is dense
in C(X; Rm), we mean density with respect to the topology of uniform convergence on
compact sets. That is, for every compact K ⊂ X, every g ∈ C(K; Rm), and every ε > 0,
there exists f ∈ F such that

Equivalently, for each compact K ⊂ X, the restrictions

∥g − f ∥K < ε.

{ f |K : f ∈ F }

are dense in C(K; Rm) with respect to ∥ · ∥K.

The space C(X; R) will be denoted by C(X).

Definition 2.3 (Tauber–Wiener function). A function σ : R → R is called a Tauber–
Wiener function if the linear span of the set

is dense in C([a, b]) for every closed interval [a, b] ⊂ R.

{σ(wt − θ) : w, θ ∈ R}

Functions with this property generate dense families of translations and dilations of
σ on every compact subset of the real line and play a central role in neural network
approximation theory. The terminology originates from the work of Chen and Chen [3],
where this condition is used in the study of operator approximation by neural networks.
Tauber–Wiener functions have also been exploited in several subsequent works (see, e.g.,
[7, 22]).

We denote by Sσ(X; Rm) the class of all single-hidden-layer topological feedforward
neural networks H : X → Rm constructed from continuous linear functionals in X ∗ and
the activation function σ (see Definition 2.1).

A scalar-valued version of the following theorem was proved in [6] for neural networks
on topological vector spaces possessing the Hahn–Banach extension property. Since every
locally convex space has this property, the scalar case follows from that result. In the
present paper we restrict attention to locally convex spaces, which are widely used in
analysis, although the results remain valid for topological vector spaces with the Hahn–
Banach extension property. For completeness, we include below a direct proof for the
vector-valued case.

5

第 6 页

Theorem 2.1. Let X be a locally convex topological vector space and assume that the
activation function σ is a Tauber–Wiener function. Then for every compact set K ⊂ X,
every function g ∈ C(K; Rm), and every ε > 0, there exists a topological neural network
H ∈ Sσ(X; Rm) such that

∥g − H∥K < ε.
In other words, the class Sσ(X; Rm) is dense in C(X; Rm) with respect to the topology of
uniform convergence on compact subsets of X.

Proof. Fix a compact set K ⊂ X, a function

g = (g1, . . . , gm) ∈ C(K; Rm),

and ε > 0. We construct a topological neural network H ∈ Sσ(X; Rm) such that ∥g −
H∥K < ε.

We equip Rm with the sup norm

∥x∥Rm = max
1≤r≤m

|xr|.

Since all norms on the finite-dimensional space Rm are equivalent, this choice does not
change the topology of uniform convergence on compact sets.

Then

Hence

∥g − H∥K = sup
x∈K

∥g(x) − H(x)∥Rm = sup
x∈K

max
1≤r≤m

|gr(x) − Hr(x)|.

∥g − H∥K = max
1≤r≤m

∥gr − Hr∥K,

where for scalar-valued functions on K we write ∥h∥K = supx∈K |h(x)|.

Thus it suffices to construct, for each r = 1, . . . , m, a scalar topological neural network

Hr such that

Fix r ∈ {1, . . . , m}. Consider the set

∥gr − Hr∥K < ε.

E = span{ eℓ(x) : ℓ ∈ X ∗ } ⊂ C(K).

This set is an algebra, since for ℓ1, ℓ2 ∈ X ∗ we have

eℓ1(x)eℓ2(x) = e(ℓ1+ℓ2)(x),

ℓ1 + ℓ2 ∈ X ∗,

and it contains the constant functions because e0 = 1.

Moreover, E separates points of K.

Indeed, if x, y ∈ K with x ̸= y, then by the
Hahn–Banach continuous extension theorem there exists ℓ ∈ X ∗ such that ℓ(x) ̸= ℓ(y)
(see, e.g., [19, Theorem 3.6]). Consequently,

eℓ(x) ̸= eℓ(y).

Therefore, by the Stone–Weierstrass theorem, E is dense in C(K) with respect to the
uniform norm.

6

第 7 页

Hence there exist ℓ1, . . . , ℓM ∈ X ∗ and real coefficients α1, . . . , αM such that

(cid:13)
(cid:13)
(cid:13)
(cid:13)
(cid:13)

gr −

M
(cid:88)

i=1

αieℓi(·)

(cid:13)
(cid:13)
(cid:13)
(cid:13)
(cid:13)K

<

ε
2

.

(2.1)

Since each ℓi is continuous and K is compact, the image ℓi(K) is a compact subset of

R. Hence there exists a compact interval [ai, bi] such that ℓi(K) ⊂ [ai, bi].

Because σ is a Tauber–Wiener function, the linear span of the family {σ(wt − θ) :
w, θ ∈ R} is dense in C([ai, bi]). Applying this property to the function t (cid:55)→ et, we obtain,
for each i, an integer Ni and real parameters ci,j, wi,j, θi,j such that

(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

et −

Ni(cid:88)

j=1

sup
t∈[ai,bi]

ci,jσ(wi,jt − θi,j)

(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

<

ε
2(1 + (cid:80)M
i=1 |αi|)

.

Since ℓi(K) ⊂ [ai, bi], substituting t = ℓi(x) gives

(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

sup
x∈K

eℓi(x) −

Ni(cid:88)

j=1

(cid:12)
(cid:12)
(cid:12)
ci,jσ(wi,jℓi(x) − θi,j)
(cid:12)
(cid:12)

<

ε
2(1 + (cid:80)M
i=1 |αi|)

.

(2.2)

Note that for each i, j, the map x (cid:55)→ wi,jℓi(x) is again a continuous linear functional
on X, that is, wi,jℓi ∈ X ∗. Thus each term σ(wi,jℓi(x) − θi,j) has the form σ(f (x) − θ)
with f ∈ X ∗.
Define

Hr(x) =

M
(cid:88)

Ni(cid:88)

αi

ci,jσ(cid:0)wi,jℓi(x) − θi,j

(cid:1),

x ∈ X.

Then Hr ∈ Sσ(X; R).

i=1

j=1

Using (2.2), multiplying the corresponding estimates by |αi| and summing over i =

1, . . . , M , we obtain

(cid:13)
(cid:13)
(cid:13)
(cid:13)
(cid:13)

M
(cid:88)

i=1

αieℓi(·) − Hr

(cid:13)
(cid:13)
(cid:13)
(cid:13)
(cid:13)K

<

ε
2

.

Combining this with (2.1) yields

∥gr − Hr∥K < ε.

Repeating the construction for each component r = 1, . . . , m, we obtain scalar topo-

logical neural networks H1, . . . , Hm. Define

H(x) = (H1(x), . . . , Hm(x)),

x ∈ X.

Then H ∈ Sσ(X; Rm) and

∥g − H∥K = sup
x∈K

max
1≤r≤m

|gr(x) − Hr(x)| < ε.

This completes the proof.

7

第 8 页

In the next section, topological neural networks on X play the role of the branch
component in a DeepONet. Theorem 2.1 provides the tool for approximating continuous
coefficient maps in the locally convex setting, where the available measurements are given
by continuous linear functionals.

The preceding result provides a universality principle for neural networks on locally
convex input spaces. An important example arises when the input space is the Banach
space C(K) endowed with the uniform norm. In this case, Theorem 2.1 yields the classical
operator approximation theorem of Chen and Chen [3].

Theorem 2.2 (Chen and Chen [3]). Let K be a compact metric space and let V ⊂ C(K)
be compact. Assume that the activation σ ∈ C(R) is a Tauber–Wiener function. Then
for every continuous functional f ∈ C(V ) and every ε > 0 there exist integers N, k ≥ 1,
points x1, . . . , xk ∈ K, and real parameters ci, θi, ξij (i = 1, . . . , N , j = 1, . . . , k) such that

(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

f (u) −

N
(cid:88)

i=1

ci σ

(cid:32) k

(cid:88)

j=1

ξij u(xj) − θi

(cid:33)(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

< ε

holds for all u ∈ V .

Proof. We regard C(K) as a Banach space and hence as a locally convex space. Thus,
Theorem 2.1 applies with X = C(K), compact set V ⊂ X, and m = 1. Therefore, there
exist N ≥ 1, continuous linear functionals ℓ1, . . . , ℓN ∈ C(K)∗, and real numbers ci, θi
such that

(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

sup
u∈V

f (u) −

N
(cid:88)

i=1

(cid:12)
(cid:12)
(cid:12)
ci σ(ℓi(u) − θi)
(cid:12)
(cid:12)

<

ε
2

.

(2.3)

Each ℓi is a continuous linear functional on C(K) and hence, by the Riesz represen-

tation theorem, admits a representation

ℓi(u) =

(cid:90)

K

u dµi

for a finite signed Borel measure µi on K.

Since V ⊂ C(K) is compact, it is equicontinuous by the Arzel`a–Ascoli theorem.
Consequently, for every δ > 0, each functional ℓi can be uniformly approximated on V by
a Riemann-type sum for the integral (cid:82)
K u dµi, i.e., by a finite linear combination of point
evaluations. That is, there exist points x1, . . . , xk ∈ K and coefficients ξij such that

(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

sup
u∈V

ℓi(u) −

k
(cid:88)

j=1

(cid:12)
(cid:12)
(cid:12)
ξij u(xj)
(cid:12)
(cid:12)

< δ,

i = 1, . . . , N.

Since σ is uniformly continuous on compact intervals containing the ranges of the

arguments, we may choose δ sufficiently small so that

(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

N
(cid:88)

i=1

sup
u∈V

ci σ(ℓi(u) − θi) −

N
(cid:88)

i=1

ci σ

(cid:32) k

(cid:88)

j=1

ξiju(xj) − θi

(cid:33)(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

<

ε
2

.

(2.4)

Combining (2.3) and (2.4) by the triangle inequality gives the desired ε-approximation.

8

第 9 页

3 DeepONets on locally convex spaces and operator

approximation

Let X be a locally convex topological vector space, V ⊂ X a compact set, and K ⊂ Rd
a compact set. We consider continuous operators

G : V −→ C(K; Rm),

where C(K; Rm) is equipped with the uniform norm ∥ · ∥K. In the DeepONet framework,
one seeks to approximate the mapping

(u, y) (cid:55)−→ G(u)(y),

(u, y) ∈ V × K,

by a structured network that separates the dependence on the input u from the depen-
dence on the coordinate y.

Definition 3.1 (Topological DeepONet). Fix integers p ≥ 1 and m ≥ 1. A topological
DeepONet consists of

1) a branch network

B : X → Rm×p,

which takes an input u ∈ X and outputs a matrix

B(u) = [b1(u) · · · bp(u)],

where each column bk : X → Rm is a topological neural network on X constructed
using continuous linear functionals from the dual space X ∗ and an activation func-
tion σ (see Definition 2.1);

2) a trunk network

T : Rd → Rp,

which takes a point y ∈ Rd as input and outputs the vector

T (y) = (t1(y), . . . , tp(y))T ,

represented by a Euclidean neural network (for example, a single-hidden-layer net-
work), whose components tk : Rd → R, k = 1, . . . , p, denote the p outputs of the
network.

The resulting DeepONet defines an operator

(cid:98)G : X → C(Rd, Rm)

given by

Equivalently,

( (cid:98)G(u))(y) = B(u) T (y),

u ∈ X, y ∈ Rd.

(cid:98)G(u)(y) =

p
(cid:88)

k=1

bk(u) tk(y).

(3.1)

The architecture of the topological DeepONet is illustrated in Figure 1.

9

第 10 页

u ∈ X
(locally convex space)

f1(u), f2(u), . . . , fr(u)
fj ∈ X ∗

y ∈ Rd

Branch network

Trunk network

[b1(u), . . . , bp(u)]

×

[t1(y), . . . , tp(y)]T

(cid:98)G(u)(y) =

p
(cid:88)

k=1

bk(u)tk(y)

Figure 1: Topological DeepONet architecture. The branch network encodes the input
element u ∈ X, where X is a locally convex space, through finitely many linear mea-
surements f1(u), . . . , fr(u) with fj ∈ X ∗. The trunk network takes y ∈ Rd as input and
yields [t1(y), . . . , tp(y)]T ∈ Rp. The outputs of the branch and trunk networks are then
multiplied to produce the final output. If X is the space of continuous functions and the
functionals fj are point evaluation functionals, fj(u) = u(xj), then the classical Deep-
ONet architecture is recovered.

10

第 11 页

Remark 3.1. When m = 1, (3.1) reduces to the classical dot-product form

(cid:98)G(u)(y) = ⟨b(u), t(y)⟩,

with b(u) ∈ Rp and t(y) ∈ Rp. This is the standard DeepONet architecture introduced
in [14]. The above definition extends this construction to inputs from general locally
convex spaces and to vector-valued outputs, while preserving the explicit branch–trunk
separation.

For example, when m = 1 and the trunk network is realized by a single-hidden-layer
neural network, the representation (3.1) can be written explicitly in terms of activation
functions. In particular, if the branch network uses linear measurements fj(u), it takes
the form

p
(cid:88)

n
(cid:88)

(cid:32) r

(cid:88)

(cid:98)G(u)(y) =

cki σ

ξkijfj(u) − θki

k=1

i=1
(cid:124)

j=1

(cid:123)(cid:122)
branch

(cid:33)

σ(ωk · y + ζk)
(cid:125)
(cid:123)(cid:122)
(cid:124)
trunk

.

(cid:125)

This coincides with the displayed formula in [14] when X is the space of continuous
functions and the functionals fj(u) are taken as point evaluations fj(u) = u(xj).

Theorem 2.1 provides approximation capability on the branch network side. On the
trunk network side, we rely on the classical density of ridge networks (i.e., single-hidden-
layer neural networks) on compact subsets of Rd. It is well known that if the activation
σ is a Tauber–Wiener function, then finite linear combinations of ridge functions

y (cid:55)−→ σ(ω · y + ζ),

y, ω ∈ Rd, ζ ∈ R,

are dense in C(K) for every compact K ⊂ Rd (see, e.g., [12, 17]). For background on
ridge functions, see [5, 18].

We now prove a universal approximation theorem for continuous operators G : V →
C(K; Rm) by finite separable expansions of the form (3.1), where the coefficient maps are
realized by branch topological neural networks on X.

Theorem 3.1. Let X be a locally convex topological vector space and let V ⊂ X be
compact. Let K ⊂ Rd be compact and let G : V → C(K; Rm) be continuous. Assume
that the activation σ ∈ C(R) is a Tauber–Wiener function.

Then for every ε > 0 there exist an integer N ≥ 1, ridge functions

ϕk(y) = σ(ωk · y + ζk),

k = 1, . . . , N,

and topological neural networks

ak : X → Rm,

k = 1, . . . , N,

constructed according to Definition 2.1, such that

sup
u∈V

sup
y∈K

(cid:13)
(cid:13)
(cid:13)
(cid:13)
(cid:13)

G(u)(y) −

N
(cid:88)

k=1

ak(u) ϕk(y)

(cid:13)
(cid:13)
(cid:13)
(cid:13)
(cid:13)Rm

< ε.

Proof. Define

W := G(V ) ⊂ C(K; Rm).

11

第 12 页

Since V is compact and G is continuous, W is compact in C(K; Rm) equipped with the
uniform norm

We equip Rm with the sup-norm

∥h∥K = sup
y∈K

∥h(y)∥Rm.

∥x∥Rm = max
1≤r≤m

|xr|.

Since Rm is finite-dimensional, all norms on Rm are equivalent; hence this choice does not
affect compactness or approximation properties, and simplifies componentwise estimates
below.

Fix an arbitrary ε > 0 and set δ := ε/4.
Let h = (h1, . . . , hm) ∈ W . For each component r = 1, . . . , m, the density of ridge

networks on K yields a finite ridge expansion

Rh,r(y) =

N (h,r)
(cid:88)

j=1

αh,r,j σ(ωh,r,j · y + ζh,r,j),

y ∈ K,

such that ∥hr − Rh,r∥K < δ.

Define the vector-valued function

Rh(y) := (Rh,1(y), . . . , Rh,m(y)),

y ∈ K.

Then

Define an open neighborhood of h in W by

∥h − Rh∥K < δ.

Uh := {h′ ∈ W : ∥h′ − h∥K < δ}.

Then for every h′ ∈ Uh,

∥h′ − Rh∥K ≤ ∥h′ − h∥K + ∥h − Rh∥K < δ + δ = 2δ.

(3.2)

Since W is compact, the open cover {Uh : h ∈ W } admits a finite subcover. Hence

there exist h(1), . . . , h(M ) ∈ W such that

W ⊂

M
(cid:91)

j=1

Uh(j).

(3.3)

Now collect all ridge functions appearing in the finitely many approximants Rh(j).

Denote the elements of this finite set by

ϕk(y) = σ(ωk · y + ζk),

k = 1, . . . , N.

By adding zero coefficients if necessary, for each j = 1, . . . , M we may write

Rh(j)(y) =

N
(cid:88)

k=1

Aj,k ϕk(y),

y ∈ K,

(3.4)

with vectors Aj,k ∈ Rm.

12

第 13 页

Since W is a compact metric space, it is paracompact. By the partition of unity
theorem (see Dugundji [4, p. 170, Theorem 4.2]), the finite open cover (3.3) admits a
continuous partition of unity subordinate to it. Hence there exist continuous functions

ηj : W → [0, 1],

j = 1, . . . , M,

such that

1) (cid:80)M

j=1 ηj(h) = 1 for all h ∈ W ;

2) supp(ηj) ⊂ Uh(j).

Define continuous coefficient maps ck : W → Rm by

ck(h) :=

M
(cid:88)

j=1

ηj(h) Aj,k,

k = 1, . . . , N.

Define

A : W → C(K; Rm),

(A(h))(y) :=

N
(cid:88)

k=1

ck(h) ϕk(y).

We claim that

∥h − A(h)∥K ≤ 2δ

for all h ∈ W.

(3.5)

Indeed, using (3.4) and the definition of ck,

A(h)(y) =

M
(cid:88)

j=1

ηj(h) Rh(j)(y),

hence

h(y) − A(h)(y) =

M
(cid:88)

j=1

ηj(h)(cid:0)h(y) − Rh(j)(y)(cid:1).

If ηj(h) ̸= 0, then h ∈ Uh(j), and therefore by (3.2) we have ∥h − Rh(j)∥K < 2δ. Thus

∥h(y) − A(h)(y)∥Rm ≤

M
(cid:88)

j=1

ηj(h) · 2δ = 2δ.

Taking the maximum over y ∈ K yields (3.5).

For each k and r = 1, . . . , m, define

bk,r(u) := (cid:0)ck(G(u))(cid:1)

r,

u ∈ V.

Each bk,r is continuous on V .

Set

Mk := max
y∈K

|ϕk(y)| < ∞.

By Theorem 2.1, for each (k, r) there exists a scalar topological network ak,r : X → R
such that

δ
N (Mk + 1)m

.

(3.6)

|bk,r(u) − ak,r(u)| <

sup
u∈V

13

第 14 页

Define

and set

ak(u) = (ak,1(u), . . . , ak,m(u)),

k = 1, . . . , N,

(cid:101)G(u)(y) :=

N
(cid:88)

k=1

ak(u) ϕk(y).

Fix u ∈ V and y ∈ K. By (3.5),

∥G(u)(y) − A(G(u))(y)∥Rm ≤ 2δ.

Moreover,

A(G(u))(y) − (cid:101)G(u)(y) =

N
(cid:88)

k=1

(ck(G(u)) − ak(u)) ϕk(y).

Using (3.6) and |ϕk(y)| ≤ Mk, we obtain

∥A(G(u))(y) − (cid:101)G(u)(y)∥Rm ≤

N
(cid:88)

k=1

δ
N (Mk + 1)m

· Mk ≤ δ.

Hence

∥G(u)(y) − (cid:101)G(u)(y)∥Rm ≤ 2δ + δ < ε.
Taking supremum over u ∈ V and y ∈ K completes the proof.

We now state an approximation theorem in the style of the dot-product DeepONet

formulation (compare [14, Theorem 2]).

Theorem 3.2. Under the assumptions of Theorem 3.1, for every ε > 0 there exists an
integer p ≥ 1, ridge functions tk(y) = σ(ωk · y + ζk) on K ⊂ Rd, and a trunk map

T (y) = (cid:0)t1(y), . . . , tp(y)(cid:1)T ∈ Rp,

y ∈ K,

together with a branch map B : X → Rm×p, whose columns are topological neural networks
on X constructed as in Definition 2.1, such that

sup
u∈V

sup
y∈K

∥G(u)(y) − B(u)T (y)∥Rm < ε.

In other words, the operator G admits an ε–approximation on V × K by a DeepONet in
the sense of Definition 3.1.

Proof. By Theorem 3.1, there exists a separable expansion

G(u)(y) ≈

p
(cid:88)

k=1

ak(u) tk(y),

tk(y) = σ(ωk · y + ζk).

Define T (y) = (t1(y), . . . , tp(y))T and define B(u) as the m × p matrix whose kth column
is ak(u). Then

p
(cid:88)

ak(u) tk(y) = B(u)T (y),

and the claimed uniform error bound is exactly the estimate obtained in Theorem 3.1.

k=1

14

第 15 页

Remark 3.2. When m = 1, the matrix–vector product reduces to a dot product:
B(u)T (y) = ⟨b(u), t(y)⟩, which is the dot-product formulation of DeepONets in [14].
Theorem 3.2 therefore extends the DeepONet approximation theorem to operators whose
input lies in a compact subset of a locally convex space and whose output functions are
Rm-valued.

Remark 3.3. In classical DeepONets, the branch network typically receives measure-
ments of the input function at finitely many sensor locations, that is, values of the form
u(x1), . . . , u(xr). Such measurements are meaningful because the input belongs to a
function space, where pointwise evaluation is naturally defined.

In the present setting, the input u is an element of a locally convex space X and need
not be a function. Consequently, pointwise sampling is not available in general. Instead,
the branch encoder accesses u through finitely many continuous linear functionals ℓ(u),
where ℓ ∈ X ∗. These functionals play the role of generalized sensors: each network uses
finitely many of them to form a finite measurement vector.

When X happens to be a function space, classical sensors are recovered as a special
case. For example, if X = C(Ω) is the space of continuous functions on a compact set
Ω endowed with the uniform norm topology, then the point evaluation maps u (cid:55)→ u(x)
are continuous linear functionals, hence belong to X ∗. Thus continuous linear functionals
provide a natural and flexible abstract measurement interface for operator learning in the
locally convex setting.

Remark 3.4. The trunk side in Theorem 3.2 is presented in ridge form because it is
one of the simplest universal approximators on compacts of Rd. In applications, one may
replace the ridge trunk by a deep neural network (ResNet, CNN, etc.) provided that
the chosen trunk class is universal on K. Likewise, on the branch side one may choose
shallow or deep topological networks on X constructed as in Definitions 2.1 and 2.2. The
approximation mechanism requires only density on compact sets, as captured abstractly
by Theorem 2.1.

The preceding remarks clarify the interpretation of the branch–trunk construction
and its relation to classical DeepONets. We now record two direct consequences of The-
orem 3.1 showing that, when the input space is a continuous function space and the ad-
missible measurements are chosen accordingly, one recovers the operator-approximation
results that form the theoretical foundation of DeepONets.

In particular, the classical Chen–Chen operator approximation theorem and the dot-
product DeepONet approximation theorem of Lu et al. appear as special cases of the
present locally convex framework.

Corollary 3.1 (see Theorem 5 in [3] or Theorem 1 in [14]). Suppose E is a Banach
space, K1 ⊂ E and K2 ⊂ Rd are compact sets in E and Rd, respectively, V ⊂ C(K1)
is compact, and G : V → C(K2) is a continuous nonlinear operator. Assume that the
activation σ ∈ C(R) is a Tauber–Wiener function.

Then for every ε > 0 there exist integers n, p, r ≥ 1, points x1, . . . , xr ∈ K1, parame-

ters cki, θki, ξkij, ζk ∈ R, and vectors ωk ∈ Rd such that

(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

G(u)(y) −

p
(cid:88)

n
(cid:88)

(cid:32) r

(cid:88)

cki σ

k=1

i=1

j=1

ξkiju(xj) − θki

(cid:33)

(cid:12)
(cid:12)
(cid:12)
σ(ωk · y + ζk)
(cid:12)
(cid:12)

< ε

for all u ∈ V and y ∈ K2.

15

第 16 页

Proof. Apply Theorem 3.1 with X = C(K1), compact V ⊂ C(K1), K = K2, and m = 1.
Then there exist p ≥ 1, ridge functions ϕk(y) = σ(ωk · y + ζk), and topological neural
networks

such that

Set

ak : C(K1) → R,

k = 1, . . . , p,

sup
u∈V

sup
y∈K2

(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

G(u)(y) −

p
(cid:88)

k=1

(cid:12)
(cid:12)
(cid:12)
ak(u) ϕk(y)
(cid:12)
(cid:12)

<

ε
2

.

M := max
1≤k≤p

sup
y∈K2

|ϕk(y)| < ∞,

η :=

ε
2p(M + 1)

.

(3.7)

For each k, apply Theorem 2.2 to the continuous functional ak|V : V → R to obtain an
approximation of the form

(cid:101)ak(u) =

n
(cid:88)

i=1

cki σ

(cid:32) r

(cid:88)

j=1

(cid:33)

ξkiju(xj) − θki

(with a common choice of r, n after unifying finitely many sensor points and adding zero
coefficients if necessary) satisfying

sup
u∈V

|ak(u) − (cid:101)ak(u)| < η,

k = 1, . . . , p.

(3.8)

Define

(cid:101)G(u)(y) :=

p
(cid:88)

k=1

(cid:101)ak(u) ϕk(y) =

p
(cid:88)

n
(cid:88)

(cid:32) r

(cid:88)

cki σ

k=1

i=1

j=1

(cid:33)

ξkiju(xj) − θki

σ(ωk · y + ζk).

Then for u ∈ V and y ∈ K2, by (3.7)–(3.8),

|G(u)(y) − (cid:101)G(u)(y)| ≤

<

G(u)(y) −

+ pηM ≤

+ pη(M + 1) = ε.

(cid:12)
(cid:12)
(cid:12)
ak(u)ϕk(y)
(cid:12)
(cid:12)

+

p
(cid:88)

k=1

|ak(u) − (cid:101)ak(u)| |ϕk(y)|

p
(cid:88)

k=1
ε
2

(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)
ε
2

Taking the supremum over u ∈ V and y ∈ K2 yields the required uniform estimate.
Corollary 3.2 (see Theorem 2 in [14]). Let E be a Banach space, K1 ⊂ E and K2 ⊂ Rd
compact sets, and let V ⊂ C(K1) be compact. Assume that G : V → C(K2) is a
continuous nonlinear operator and that the activation σ ∈ C(R) is a Tauber–Wiener
function.

Then for every ε > 0 there exist integers r, p ≥ 1, points x1, . . . , xr ∈ K1, and

continuous mappings

such that

B : Rr → Rp,

T : Rd → Rp,

(cid:12)G(u)(y) − (cid:10)B(cid:0)u(x1), . . . , u(xr)(cid:1), T (y)(cid:11)(cid:12)
(cid:12)

(cid:12) < ε

for all u ∈ V and y ∈ K2.

Moreover, the maps B and T may be chosen from any class of neural networks on Rr
and Rd, respectively, that is dense in the corresponding spaces of continuous functions on
compact sets (e.g. fully connected, residual, or convolutional architectures).

16

第 17 页

Proof. Fix ε > 0. By the preceding corollary, there exist integers r, n, p ≥ 1, points
x1, . . . , xr ∈ K1, real parameters cki, θki, ξkij, ζk, and vectors ωk ∈ Rd such that

sup
u∈V

sup
y∈K2

(cid:12)
(cid:12)
(cid:12)
(cid:12)
(cid:12)

G(u)(y) −

p
(cid:88)

n
(cid:88)

(cid:32) r

(cid:88)

cki σ

k=1

i=1

j=1

ξkiju(xj) − θki

(cid:33)

(cid:12)
(cid:12)
(cid:12)
σ(ωk · y + ζk)
(cid:12)
(cid:12)

< ε.

(3.9)

Define B : Rr → Rp by

B(z1, . . . , zr) :=

(cid:32) n

(cid:88)

(cid:32) r

(cid:88)

c1i σ

i=1

j=1

(cid:33)

ξ1ijzj − θ1i

, . . . ,

n
(cid:88)

i=1

cpi σ

(cid:32) r

(cid:88)

j=1

ξpijzj − θpi

(cid:33) (cid:33)
,

and define T : Rd → Rp by

T (y) := (cid:0)σ(ω1 · y + ζ1), . . . , σ(ωp · y + ζp)(cid:1).

Both maps are continuous because σ is continuous and they are finite linear combinations
and compositions of continuous functions.

For u ∈ V and y ∈ K2, substituting z = (u(x1), . . . , u(xr)) gives

(cid:10)B(u(x1), . . . , u(xr)), T (y)(cid:11) =

p
(cid:88)

n
(cid:88)

(cid:32) r

(cid:88)

cki σ

k=1

i=1

j=1

(cid:33)

ξkiju(xj) − θki

σ(ωk · y + ζk).

Combining this identity with (3.9) yields

sup
u∈V

sup
y∈K2

(cid:12)G(u)(y) − (cid:10)B(cid:0)u(x1), . . . , u(xr)(cid:1), T (y)(cid:11)(cid:12)
(cid:12)

(cid:12) < ε,

as required.

The final statement follows by approximating the continuous maps B and T uniformly

on the relevant compact sets by networks from any dense architecture class.

Remark 3.5. Lanthaler et al. [11] establish a universality result for DeepONets in a
probabilistic setting: the input space is equipped with a probability measure and the
approximation is formulated in an L2 metric. Within this framework they obtain approx-
imation results for measurable operators, thereby removing the continuity and compact-
ness assumptions that appear in the classical operator approximation theorem of Chen
and Chen [3]. Their work also provides quantitative error bounds and complexity esti-
mates for DeepONets, which lie outside the scope of the present approximation-theoretic
framework.

In comparison with the Chen–Chen theorem, the universality result of Lanthaler et al.
comes at the expense of considering a weaker distance, namely L2 instead of the uniform
L∞ distance (see [11, Remark 3.1]).

Thus, while [11] relaxes continuity and compactness assumptions by working in a prob-
abilistic L2 framework, the present work preserves uniform approximation and extends
operator approximation theory beyond Banach input domains to general locally convex
spaces. In particular, the classical Chen–Chen theorem and the DeepONet approximation
theorem of Lu et al. appear as special cases of our framework.

17

第 18 页

4 Examples illustrating Theorem 3.1

In this section we illustrate Theorem 3.1 for several important normed and locally convex
spaces. In all examples the admissible measurements are continuous linear functionals
from the dual space X ∗. By Theorem 3.1, for a compact set V ⊂ X and a continuous
operator

G : V → C(K; Rm),

K ⊂ Rd compact,

we obtain approximations of the form

G(u)(y) ≈

N
(cid:88)

k=1

ak(u) σ(ωk · y + ζk),

u ∈ V, y ∈ K,

(4.1)

where each coefficient map ak : X → Rm is a topological neural network on X constructed
as in Definition 2.1.
Example 1. Let X = Rd or, more generally, let

X = Mn×p(R)

be the vector space of all n × p real matrices equipped with any norm. Then X is a finite-
dimensional Banach space and every continuous linear functional on X can be written in
the form

A (cid:55)−→ trace(W T A),

A ∈ X,

for some matrix W ∈ Mn×p(R).

Let V ⊂ X be compact and let G : V → C(K; Rm) be continuous. Applying Theo-
rem 3.1 yields approximations of the form (4.1), where each coefficient map ak : X → Rm
is a topological neural network on X in the sense of Definition 2.1. More precisely, there
exist matrices Wk,1, . . . , Wk,rk ∈ Mn×p(R), vectors ck,i ∈ Rm, and scalars θk,i ∈ R such
that

rk(cid:88)

ak(A) =

ck,i σ(cid:0)trace(W T

k,iA) − θk,i

(cid:1),

A ∈ X.

i=1

Thus the branch part of the network uses finitely many linear measurements of the
input matrix A of the form A (cid:55)→ trace(W T A), while the trunk part consists of ridge
functions on K.

Example 2. Let 1 ≤ p < ∞ and let X = ℓp, the Banach space of real sequences
x = (x1, x2, . . . ) with

(cid:32) ∞
(cid:88)

(cid:33)1/p

∥x∥p =

|xn|p

< ∞.

Then the continuous dual is X ∗ = ℓq, q = p
ℓp can be written in the form

p−1 , and every continuous linear functional on

n=1

x (cid:55)−→

∞
(cid:88)

n=1

wnxn,

w = (w1, w2, . . . ) ∈ ℓq.

Let V ⊂ ℓp be compact and let G : V → C(K; Rm) be continuous. Applying Theo-
rem 3.1 yields approximations of the form (4.1), where each coefficient map ak : X → Rm

18

第 19 页

is a topological neural network on X in the sense of Definition 2.1. More precisely, there
exist vectors

vectors ck,i ∈ Rm, and scalars θk,i ∈ R such that

w(k,1), . . . , w(k,rk) ∈ ℓq,

ak(x) =

rk(cid:88)

i=1

ck,i σ

(cid:32) ∞
(cid:88)

n=1

(cid:33)

w(k,i)

n xn − θk,i

,

x ∈ X.

Example 3. Let X = c0, the Banach space of real sequences converging to zero, equipped
with the sup norm. Then the continuous dual is X ∗ = ℓ1, and every continuous linear
functional on c0 can be written in the form

x (cid:55)−→

∞
(cid:88)

n=1

wnxn,

w = (w1, w2, . . . ) ∈ ℓ1.

Let V ⊂ c0 be compact and let G : V → C(K; Rm) be continuous. Applying Theo-
rem 3.1 yields approximations of the form (4.1), where each coefficient map ak : X → Rm
is a topological neural network on X in the sense of Definition 2.1. More precisely, there
exist vectors

vectors ck,i ∈ Rm, and scalars θk,i ∈ R such that

w(k,1), . . . , w(k,rk) ∈ ℓ1,

ak(x) =

rk(cid:88)

i=1

ck,i σ

(cid:32) ∞
(cid:88)

n=1

(cid:33)

w(k,i)

n xn − θk,i

,

x ∈ X.

Example 4. Let (Ω, µ) be a measure space and let X = Lp(Ω, µ), 1 ≤ p < ∞. If p = 1,
assume in addition that (Ω, µ) is σ-finite. Then X is Banach and its continuous dual
is X ∗ = Lq(Ω, µ), where q = p
p−1 (with q = ∞ if p = 1), and every continuous linear
functional on X has the form
(cid:90)

f (cid:55)−→

f (x)g(x) dµ(x),

g ∈ Lq(Ω, µ).

Ω

Let V ⊂ Lp(Ω, µ) be compact and let G : V → C(K; Rm) be continuous. Applying
Theorem 3.1 yields approximations of the form (4.1), where each coefficient map ak :
X → Rm depends on finitely many integral measurements. More precisely, there exist
functions gk,1, . . . , gk,rk ∈ Lq(Ω, µ), vectors ck,i ∈ Rm, and scalars θk,i ∈ R such that

ak(f ) =

rk(cid:88)

i=1

ck,i σ

(cid:18)(cid:90)

Ω

f (x)gk,i(x) dµ(x) − θk,i

,

f ∈ X.

(cid:19)

Example 5. Let X = S(Rn), the Schwartz space of rapidly decreasing smooth func-
tions. This is a Fr´echet space whose continuous dual X ∗ = S ′(Rn) consists of tempered
distributions. Every continuous linear functional has the form

f (cid:55)−→ ⟨T, f ⟩,

T ∈ S ′(Rn).

Let V ⊂ S(Rn) be compact and let G : V → C(K; Rm) be continuous. Applying
Theorem 3.1 yields approximations of the form (4.1), where each coefficient map ak : X →

19

第 20 页

Rm depends on finitely many distributional measurements. More precisely, there exist
tempered distributions Tk,1, . . . , Tk,rk ∈ S ′(Rn), vectors ck,i ∈ Rm, and scalars θk,i ∈ R
such that

rk(cid:88)

ak(f ) =

ck,i σ(⟨Tk,i, f ⟩ − θk,i) ,

f ∈ X.

i=1

Example 6. Let X = D(U ), the space of smooth compactly supported functions on
an open set U ⊂ Rn. Its continuous dual X ∗ = D′(U ) is the space of distributions, and
every continuous linear functional has the form

f (cid:55)−→ ⟨T, f ⟩,

T ∈ D′(U ).

Let V ⊂ D(U ) be compact and let G : V → C(K; Rm) be continuous. Applying
Theorem 3.1 yields approximations of the form (4.1), where each coefficient map ak :
X → Rm depends on finitely many distributional measurements. More precisely, there
exist distributions Tk,1, . . . , Tk,rk ∈ D′(U ), vectors ck,i ∈ Rm, and scalars θk,i ∈ R such
that

rk(cid:88)

ak(f ) =

ck,i σ(⟨Tk,i, f ⟩ − θk,i) ,

f ∈ X.

i=1

These examples illustrate that Theorem 3.1 provides a unified approximation principle
for continuous operators on compact subsets of finite-dimensional spaces, classical Banach
spaces of sequences and functions, and important non-normable locally convex spaces
arising in analysis.

5 Conclusion

We present a topological extension of DeepONets in which the operator input lies in an
arbitrary locally convex topological vector space and the network architecture is con-
structed using continuous linear functionals from the dual space. Under the assumption
that the activation function is a Tauber–Wiener function, we prove a universal approx-
imation theorem showing that continuous nonlinear operators can be approximated on
compact subsets of the input space by finite separable expansions.

The classical Chen–Chen operator approximation theorem and the dot-product Deep-
ONet approximation theorem of Lu et al. arise as special cases of our results. Several
examples illustrate that the main theorem applies to a wide range of spaces, including
infinite-dimensional Banach spaces and non-normable locally convex spaces.

References

[1] L. Bhan, P. Qin, M. Krstic and Y. Shi, Neural operators for predictor feedback

control of nonlinear delay systems, arXiv preprint, arXiv:2411.18964, 2025.

[2] S. Cai, Z. Wang, L. Lu, T. A. Zaki and G. E. Karniadakis, DeepM&Mnet: Inferring
the electroconvection multiphysics fields based on operator approximation by neural
networks, J. Comput. Phys. 436 (2021), Art. no. 110296.

20

第 21 页

[3] T. Chen and H. Chen, Universal approximation to nonlinear operators by neural net-
works with arbitrary activation functions and its application to dynamical systems,
IEEE Trans. Neural Netw. 6 (1995), no. 4, 911-917.

[4] J. Dugundji, Topology, Allyn and Bacon, Inc., Boston, MA, 1966.

[5] V. E. Ismailov, Ridge functions and applications in neural networks, Mathematical
Surveys and Monographs, 263. American Mathematical Society, Providence, RI,
2021.

[6] V. E. Ismailov, Universal approximation theorem for neural networks with inputs
from a topological vector space, Information Process. Lett. 193 (2026), Paper no.
106623.

[7] V. E. Ismailov, On shallow feedforward neural networks with inputs from a topolog-

ical space, Ann. Math. Artif. Intell. (2026).

[8] Q. Jiang, M. Salvadori, D. Ota, V. Shankar and K. Shukla, Complex valued deep
operator network (DeepONet) for three dimensional Maxwell’s equations: G ∈ Cm×n,
arXiv preprint, arXiv:2411.18733, 2024.

[9] T. O. de Jong, K. Shukla and M. Lazar, Deep Operator Neural Network Model

Predictive Control, IEEE Open J. Control Syst. 4 (2025), 501-517.

[10] M. Krstic, L. Bhan and Y. Shi, Neural operators of backstepping controller and
observer gain functions for reaction–diffusion PDEs, Automatica 164 (2024), Art.
no. 111649.

[11] S. Lanthaler, S. Mishra and G. E. Karniadakis, Error estimates for DeepONets: a
deep learning framework in infinite dimensions, Trans. Math. Appl. 6 (2022), no. 1,
141 pp.

[12] M. Leshno, V. Ya. Lin, A. Pinkus and S. Schocken, Multilayer feedforward networks
with a nonpolynomial activation function can approximate any function, Neural
Netw. 6 (1993), 861-867.

[13] C. Lin, Z. Li, L. Lu, S. Cai, M. Maxey and G. E. Karniadakis, Operator learning for
predicting multiscale bubble growth dynamics, J. Chem. Phys. 154 (2021), 104118.

[14] L. Lu, P. Jin, G. Pang, Z. Zhang and G. E. Karniadakis, Learning nonlinear operators
via DeepONet based on the universal approximation theorem of operators, Nat.
Mach. Intell. 3 (2021), no. 3, 218-229.

[15] Z. Mao, L. Lu, O. Marxen, T. A. Zaki and G. E. Karniadakis, DeepM&Mnet for
hypersonics: predicting the coupled flow and finite-rate chemistry behind a nor-
mal shock using neural-network approximation of operators, J. Comput. Phys. 447
(2021), 110698.

[16] V. Oommen, K. Shukla, S. Goswami, R. Dingreville and G. E. Karniadakis, Learning
two-phase microstructure evolution using neural operators and autoencoder archi-
tectures, npj Comput. Mater. 8 (2022), no. 1, Art. no. 190.

21

第 22 页

[17] A. Pinkus, Approximation theory of the MLP model in neural networks, Acta Nu-

merica 8 (1999), 143-195.

[18] A. Pinkus, Ridge functions, Cambridge Tracts in Mathematics, 205. Cambridge Uni-

versity Press, 2015.

[19] W. Rudin, Functional analysis, Second edition. International Series in Pure and

Applied Mathematics. McGraw-Hill, Inc., New York, 1991.

[20] K. Shukla, V. Oommen, A. Peyvan, M. Penwarden, N. Plewacki, L. Bravo, A.
Ghoshal, R. M. Kirby and G. E. Karniadakis, Deep neural operators as accurate
surrogates for shape optimization, Eng. Appl. Artif. Intell. 129 (2024), Art. no.
107615.

[21] K. Shukla, J. Ratchford, L. Bravo, V. Oommen, N. Plewacki, A. Ghoshal and G.
Karniadakis, Deep operator learning-based surrogate models for aerothermodynamic
analysis of AEDC hypersonic waverider, arXiv preprint, arXiv:2405.13234, 2024.

[22] M. E. Valle, W. L. Vital and G. Vieira, Universal approximation theorem for vector-
and hypercomplex-valued neural networks, Neural Netw. 180 (2024), 106632.

[23] S. Wang, H. Wang and P. Perdikaris, Learning the solution operator of parametric
partial differential equations with physics-informed DeepONets, Sci. Adv. 7 (2021),
eabi8605.

22

第 23 页

(本页无文本内容)