Topological Vector Spaces Part 1: Topologies
To start out this new blog, I will be writing a sequence of posts on topological vector spaces, my favorite object in math. This is the first post.
Prerequisites: Real analysis, basic topology. No functional analysis is assumed.
Needless to say, there is only so much I can write in a blog post. You are highly encouraged to follow along via other texts. In addition, Wikipedia is a good point to start Wikipedia has a surprising amount of detail on certain math topics, and topological vector spaces are one of them.
Suggested Reading
- Topological Vector Spaces, Distributions and Kernels, Treves. The reference on topological vector spaces.
Motivation
Arguably the two most important structures identified in math in the last century are topological spaces in pure math and vector spaces in all of math, but more so applied math. One can combine the two structures. While focusing on purely the global topological behavior built over trivial local vector space structure, one obtains manifolds (and a slight extension are vector bundles). While focusing purely on the local topological structure held together by trivial global vector space structure, one obtains topological vector spaces.
For a more practical motivation, many applied math problems amount to solving $f(x) = 0$ where $f : X \rightarrow Y$ is generally a nonlinear continuous map between topological vector spaces. The topology is essential for treating this problem as continuous rather than discrete. The essentiality of vector space structure is more debatable, but it is certainly very convenient for many reasons, and we work with vector spaces if we can. Prototypically, $X, Y$ are real-valued function spaces. We have no general means of finding the zero exactly, so generically our solution is iterative. We would like to find a sequence $x_m$ such that $y_m := f(x_m)$ has $y_m$ converges to zero, or $y_m \rightarrow 0$ in $Y$. 1
But what does $y_m \rightarrow 0$ even mean? That is precisely the data the topology carries. In the most trivial case, $Y$ has a canonical choice of norm, there is a real-valued error $\Vt y_m \Vt_{Y}$ for free, and we just need $\Vt y_m \Vt \rightarrow 0$. In a more general setup, we may have finitely or countably many seminorms $p_k$, and each $p_k$ tracks a different kind of error that may have some physical significance. For example, if we were solving a partial differential equation, each $p_k$ might measure the error in a specific spatial, temporal, or frequency region. To have $y_m \rightarrow 0$, for each $k$, we need $p_k(y_m) \rightarrow 0$, though this may happen at different rates, and this data might be valuable.
If there are finitely many $p_k$, we can attempt to summarize them as a single norm $\Vt y \Vt := \sum_{k} p_k(y)$, or if countably many, a no longer homogenous metric
\[ d(0, y) := \sum_{k=1}^\infty 2^{-k} \frac{p_k(y)}{1 + p_k(y)} .\]But this forgets the underlying physical significance of $p_k$, and I would make the argument that it is valuable to keep the $p_k$’s separate, and study how $p_k(y_m) \rightarrow 0$ for each $k$ can happen.
Topological Vector Spaces
For convenience, we always work over the field $\R$ and only of locally convex Hausdorff topological vector spaces.
- The topology on a topological vector space is Hausdorff 2, translation invariant, and addition and scalar multiplication are continuous.
- By translation invariance, a linear map between topological vector spaces is continuous if the pre-image of a neighborhood of the origin is a neighborhood of the origin.
- A sequence $x_n$ is Cauchy if for any neighborhood of the origin $U$, there is $N$ such that $n,m > N$ implies $x_n - x_m \in U$.
- A sequence $x_n$ converges to $x$ if for any neighborhood of the origin $U$, there is $N$ such that $n > N$ implies $x_n - x \in U$.
It is typical to work with a local base at the origin. A local base $\Lambda$ is a collection of sets containing the origin which generates the topology, in the sense that any arbitrary neighborhood of the origin $U$ must contain some open set $tV \sbe U$ where $V \in \Lambda$ is in the local base, $t > 0$. We assume that the sets $V$ are chosen to be balanced and convex, balanced meaning $tV \sbe V$ for any $\vt t \vt \leq 1$. 3 A set that is both balanced and convex is called absolutely convex.
In many cases, we use neighborhoods of the origin to test that another object has properties which hold for sufficiently small open sets, and it suffices to test the local base generating the open sets. For example, for continuity of a linear map, it is enough to check that pre-images of elements in the local base is a neighborhood of the origin. For Cauchy and convergent sequences, it is also enough to test local bases.
Given an arbitrary collection $\Lambda$, if it were to generate a topological vector space, addition and multiplication must be continuous. Continuity of addition is easy, the pre-image of an absolutely convex $V$ contains $V/2 \times V/2$.
Continuity of multiplication is less obvious. For any neighborhood of the origin $U$ and non-zero $x \in X$, continuity of $t \mapsto tx$ at $0$ implies that $tx \in U$ for all $t$ in a neighborhood of $0$. 4 Geometrically, the line through $0$ and $x$ intersected with $U$ must contain a non-trivial interval. Thus, neighborhoods of the origin are fat, they have non-trivial content in all directions.
Definition. Given arbitrary subsets $U, E \sbe X$ , if there exists $s > 0$ such that $tE \sbe U$ for all $t\leq s$, we say that $U$ absorbs $E$.
We can think of $U$ absorbs $E$ meaning $U$ contains $E$ up to scalar constants. If $U$ is a neighborhood of zero and $\Lambda$ is a local base, $U$ must absorb some $V \in \Lambda$. If $U$ is balanced, then it suffices to have some $t > 0$ with $tE \sbe U$.
Continuity of multiplication implies that neighborhoods of zero must absorb singletons, and consequently finite sets. Therefore, to declare a local base $\Lambda$, in addition to making $V \in \Lambda$ absolutely convex, we make sure that $V$ absorbs singletons, so that $\Lambda$ generates a valid topology.
There is a bijection between absolutely convex, singleton-absorbing sets and seminorms. A seminorm $p : X \rightarrow [0,\infty)$ is a function that is homogenous, $p(tx) = \vt t \vt p(x)$, and satisfies the triangle inequality $p(x+y) = p(x) + p(y)$. Given $V$, the associated Minkowski functional
$$ p_V(x) = \inf \{ t > 0 : x \in tE \} $$is a seminorm, and conversely, given a seminorm $p$, the ball $B_p = \{x \in X : p(x) < 1 \}$ is an absolutely convex, singleton-absorbing set.
Exercise. Verify the above bijection, show that absolutely convex corresponds to homogenous plus triangle inequality, and singleton-absorbing corresponds to $p(x) < \infty$.
Therefore, it is common to instead specify a collection of seminorms instead of a local base. The condition $B_p$ absorbs set $E$ is equivalent to $p(E) < \infty$. Often $p$ is not a norm, meaning that there is some nonzero $x$ where $p(x) = 0$. However, we assume that all topological vector spaces are Hausdorff, meaning that the family $\Lambda$ separates any nonzero $x$ from the origin, or equivalently, for any $x$ nonzero, $p(x) > 0$ for some seminorm $p$.
Definitions.
- A normed space is a Hausdorff locally convex topological vector space whose topology is generated by a single absolutely convex, absorbing set, or equivalently, a single norm.
- A Banach space is a Cauchy-complete normed space.
- A F-space is a Hausdorff locally convex topological vector space whose topology is generated by countably many absolutely convex, absorbing sets, or equivalently, countably many seminorms.
- A Frechet space is a Cauchy-complete F-space.
The significance of an F-space with countably many seminorms $p_k$ is that it is metrizable with the translation-invariant metric given in the motivations section. Metrizable gives many topological properties for free, for example, sequential compactness is equivalent to compactness. But of course, the specific choice of metric is arbitrary. The converse also holds, if a topological vector space is metrizable, then countably many balls $B(0,1/k)$ form a local base. The balls need not be homogenous with respect to the radius, so countably many is necessary.
Motivating Examples
Let us now study some motivating examples of topological vector spaces. These examples are analogous to a fairly wide class of function spaces one frequently sees in analysis.
- The space of real-valued sequences $\R^\N$ with a countable family of seminorns $p_k(x_n) = \sup_{n \leq k} \vt x_n \vt$.
- The space of bounded sequences $\ell^\infty$ with norm $\Vt x_n \Vt_{\ell^\infty} = \sup_{n \in \N} \vt x_n \vt$.
- The space of sequences that vanish at infinity $c_0$, a subspace of $\ell^\infty$ with the same norm.
- The Schwartz space of rapidly decaying sequences $\mathcal{s}$. The sequence $x_n$ belongs to $\mathcal{s}$ if and only if $x_n$ decays to zero faster than any polynomial. $\mathcal{s}$ has a countable family of seminorms $$ p_k(x_n) = \sup_{n \in \N} \vt n^k x_n \vt. $$
- The space of eventually zero sequences $c_{0,0}$. A local base of $c_{0,0}$ is $$ V_{y_n} = \{ x_n \in c_{0,0} : \vt x_n \vt < y_n \text{ for all } n \in \N \}$$as $y_n$ ranges over all strictly positive sequences.
Exercise. Verify that the seminorms and local bases defined above are singleton-absorbing. Moreover, verify that all five spaces above are complete. Hint: 5
The reason why we place the different topologies on the different spaces is that we want the topology to play well with the underlying space. For applications we usually want the space to be complete. The $\ell^\infty$ norm can be placed on $\mathcal{s}$ and $c_{0,0}$, but in these cases the topology is too weak, there are too many Cauchy sequences that do not converge, hence we pick the stronger topologies given.
Conversely, sometimes we may first think of the seminorms, and define the space to be all elements in a larger space which are controlled by the seminorms. In the case of the Schwartz space $\mathcal{s}$, I haven’t explicitly defined what decay to zero faster than any polynomial is, but it means $x_n \in \R^\N$ such that $p_k(x_n) < \infty$ for all the associated seminorms.
The Banach spaces $\ell^\infty$ and $c_0$ do not require much explanation. Our two Frechet spaces $\R^\N$ and $\mathcal{s}$ have the property that $p_{k+1} \geq p_k$ is an increasing sequence of seminorms, or in terms of the local base $V_{k+1} \sbe V_k$ is decreasing.
Exercise. For $\R^\N$ and $\mathcal{s}$, for any $k \in \N$, exhibit some $x_n$ such that $p_k(x_n) = 1$ but $p_{k+1}(x_n)$ is arbitrarily large. In view of this, a sequence $x_n^m$ of points indexed by $m$ can converge with respect to a single seminorm $p_k$ but not with respect to all seminorms or the full topology.
The space $c_{0,0}$ is perhaps the most interesting and least intuitive at first glance. We can think of $c_{0,0} = \bigcup_{i=0}^\infty X_i$, where $X_i$ is the subspace of sequences which are zero after index $i$. Each $X_i$ is finite dimensional, and it only has one non-trivial topology, the one generated by the supremum norm.
Exercise. Show that $U \sbe c_{0,0}$ is a neighborhood of the origin if and only if $U \cap X_i$ is a neighborhood of the origin for all $i$.
To emphasize, if we placed another topology on $c_{0,0}$, for example the topology of $\R^\N$, $\ell^\infty$, or $\mathcal{s}$, then it is true that $U$ is neighborhood implies $U \cap X_i$ is a neighborhood, but the converse does not hold. A counter-example is the set $V_{y_n}$ where $y_n = 2^{-n}$, which is too thin to absorb any open set from the local base of the other topologies.
Definition. Suppose that vector space $X$ is a union of increasing subspaces $X = \bigcup_{i=1}^\infty X_i$, each $X_i$ has a Frechet (Banach) topology such that $X_i$ has the subspace topology with respect to $X_j$ for $i < j$. Then, the LF (LB) topology on $X$ is the strongest locally convex topology such that all $X_i \hookrightarrow X$ is continuous.
When $X$ has the LF topology, $U$ is neighborhood of the origin if and only if $U \cap X_i$ is a neighborhood of the origin in $X_i$. A concrete local base of $X$ is $\bigcup_{i=1}^\infty V_i$ where each $V_i \sbe X_i$ is an element of the local base of $X_i$.
We make a comparison to continuous function spaces. Note that if $\Omega \sbe \R^d$ is an open domain, then there is an increasing sequence of compact sets $K_i \sbe \Omega$ that cover $\Omega$. In other words, $\Omega$ is $\sigma$-compact.
- The analog of $\R^\N$ is $C(\Omega)$, the space of continuous functions on an open domain with the topology of uniform convergence on compacts. Explicitly, seminorms are $p_k(f) = \sup_{x \in K_k} \vt f(x) \vt$.
- The analog of $\ell^\infty$ and $c_0$ are the space of bounded continuous functions $C_b(\Omega)$ and the space of continuous functions vanishing at infinity $C_0(\Omega)$, respectively, with the usual supremum norm.
- The analog of $c_{0,0}$ is the space of compactly supported functions on an open domain $C_c(\Omega) = \bigcup_{i=1}^\infty C_c(K_i)$. The topology on $C_c(\Omega)$ is the final topology with respect to the inclusions $C_c(K_i) \hookrightarrow C_c(\Omega)$.
The space $\mathcal{s}$ doesn’t have as precise an analog, but if we think of $p_k$ has having increasing control over the decay of $x_n$, then a similar case is the space of smooth functions $C^\infty([0,1])$ whose seminorms $p_k = \sup_{x \in [0,1]} \vt f^{(i)} (x) \vt $ have increasing control over the smoothness of a function $f$. 6 7
We can further push the comparisons. One important example is the space of test functions $C_c^\infty(\Omega) = \bigcup_{i=1}^\infty C_c^\infty(K_i)$, each $C_c^\infty(K_i)$ has a Frechet space topology like $\mathcal{s}$, and the global $C_c^\infty(\Omega)$ has a inductive limit topology like $c_{0,0}$.
-
Of course, not all topological spaces are sequential. This adds a further layer of complexity to the problem, and certainly it makes an interesting question of what to do in practice, since we can only ever compute a finite sequence. Nevertheless, my intent is not to push the discussion most general. I find the countably many seminorms and Frechet space setup illustrative enough for here. ↩︎
-
Yes, we certainly will assume Hausdorff. One can in fact prove that a topological group or vector space that is Hausdorff is automatically completely regular, a closed set and a point can be separated by a continuous function. This is required by continuity of addition. ↩︎
-
We lose no generality by assuming that sets are balanced. If the topology generated by a local base $\Lambda$, we can always refine the base to contain balanced sets. Convexity restricts our topological vector spaces to locally convex ones, and this makes various parts of the theory cleaner. In practice most spaces we encounter are locally convex. The one notable exception is $L^0$, the space of random variables under the topology of convergence in measure. ↩︎
-
This statement only says that multiplication is continuous in $t$, not continuous in both $t, x$ jointly. Generally, it is not true that separately continuous linear maps are jointly continuous (we will study this more in a future post). Nevertheless, verify that in our $\R \times X \rightarrow X$ setup, continuity in $t$ is enough to imply continuity in both $t, x$. ↩︎
-
Since we are working with sequence spaces, it should be easy. Given a Cauchy sequence indexed by $m$ of sequences $x^m_n$, for each $n \in \N$, defined $x_n$ and verify that this step is well-defined. Finally, check that the overall sequence $x_n$ belongs to the space. If you get stuck on the last example, you might find it helpful to see the next post. ↩︎
-
The seminorms are not strictly increasing, since polynomials of degree $k$ have zero $p_k$ seminorm yet non-zero lower order seminorms. Often we instead take the full norm $\Vt f \Vt_{C^k} = \sum_{i=0}^k p_k(f)$ which is strictly increasing and has control over all previous seminorms. In any case, the entire countable collection of norms is necessary to generate the right topology. ↩︎
-
There is also the Schwartz space $\mathcal{S}$ of smooth functions from $\R^n$ to $\C$ whose derivatives decay faster than any polynomial. The seminorms are
$$ p_{\alpha, \beta}(f) = \sup_{x \in \R^n} \vt x^\beta (\partial^\alpha f) (x) \vt $$as $\alpha, \beta$ ranges over all (countably many) multi-indices, which controls both decay and smoothness. The Fourier transform precisely says that decay and smoothness are interchangable. ↩︎