About this researchFixed report · 743 generations · 4 Sep 2026
Question

Do emerge_v2 narratives move through semantic space, or does the system keep rephrasing the same region?

Method

Each artistic_statement is encoded with text-embedding-3-large. Distances show relative semantic displacement; the chart boundary is the edge of the observed corpus, not the limit of possible themes.

Finding

The report separates global drift, local variation and returns. It prevents unusual wording from being mistaken for a new artistic direction.

How to read it

Read the charts in sequence: corpus scale, trajectory, local transitions, then text examples at boundaries and in dense regions.

Drift of emerge_v2 plots — mathematics and semantics of displacement

Drift of emerge_v2 plots — mathematics and semantics of displacement · version 2, full corpus

743 generations, 1458 works (2026-06-29 02:29 → 2026-08-01 09:56) · “plot” = artistic_statement of the work · text-embedding-3-large embeddings, 3072d, raw space · fresh read-only prod DB dump, with no new calls to generative models. Hover over any point in any graph — the title and plot summary will pop up.

What the data showed

0 · How this works — from meaning to vector and back

The entire report rests on one technique: text is turned into a point in a multidimensional space, and then geometry works with meaning. Below are nine screens: the first explains the technique itself with a simple picture, the others break it down step by step on the real data of the corpus — the same vectors, the same works, the same numbers as in all other sections. Click the numbers.
step 1 · intuition
What does it even mean to “turn meaning into a vector”

Before calculating — a simple picture, without any corpus data. The idea of embedding is one: map each meaning to a point in space so that similar meanings stand near each other, and distant ones far apart. Then “similar in meaning” becomes simply “nearby in space”, and text can be worked with using a ruler. Below is a cloud of ordinary words rotating in three dimensions: semantic neighbors stay together.

The cloud slowly rotates — three axes are shown for clarity; each point is a word, with depth conveyed by size and brightness. Groups: animals, transport, weather, feelings. Hover to stop rotation.

Why there are not three axes, but thousands

In three dimensions, only a few neighbors fit nearby. “Cat” is close both to “dog” (animal), and to “warm”, and to “soft”, and to “domestic”, but three axes are not enough to fit all these kinds of closeness at the same time without stepping on each other.

That is why the real model takes not three axes, but 3072. In such a space there is enough room: every shade of meaning gets its own direction, and closeness stops being accidental. It is impossible to draw 3072 axes, so here there are three, and further on numbers, not the picture, take over.

Classic check of the idea: if meanings have directions, then “king” − “man” + “woman” leads approximately to the point “queen”. Meaning has become arithmetic — this is vectorization.
step 2 · vectorization
Text becomes a number

Now for real. The embedding model reads the plot of the work — text on the right — and outputs 3072 numbers; pattern on the left and these are those numbers, laid out in cells. None of them means anything on its own: the meaning is recorded in the pattern as a whole. The model is trained so that texts about similar things receive similar patterns — this is the only property on which the entire report rests. Below are the actual vectors of three works from the corpus, all 3072 coordinates.

Hover over a cell to see the coordinate number and its value.

Right to left

Panel below — original plot of the work, the very text that went into the model. The pattern on the left is its result: the same 3072 numbers, laid out in a grid of 64 per row only so that it fits on the screen. Change the work with the buttons — both the text and the pattern will change: a different meaning is a different point in space.

What is visible on the coordinate map

Blue means negative values, orange means positive values, dark means near zero. The pattern looks like noise, and that is normal: an embedding has no “color axis” or “emotion axis”; features are smeared across all coordinates at once.

Switch the work — the pattern will change entirely, but comparing them by eye is useless. That is exactly why what comes next is not the eye, but the dot product.

vector length—
coordinates are noticeably ≠ 0—
modeltext-embedding-3-large
step 3 · metric
What “close” means

Similarity of two plots is cosine of the angle between their vectors: multiply the coordinates pairwise and add them up. 1 means the same direction, 0 means no connection. The number says nothing by itself; it has to be calibrated against the corpus: here is the distribution of all 247 456 pairs of works and four reference points on it.

Calibration

Two random works in the corpus resemble —. Temporally adjacent ones are — at 0.65. The gap is tiny: 80% of the distance between random works is already gained in one step.

Hence the main conclusion of the report: neighboring works are almost as different as unrelated ones. This is not a malfunction; it is a measure of how widely the engine scatters plots around one theme.

step 4 · dimensionality
3072 dimensions, of which ~64 work

The space is enormous, but the cloud of plots occupies a thin pancake within it. On the left is how much variance falls on each direction after rotating the axes toward the principal components; on the right is the cumulative share. The first 10 directions hold 31% of the total spread, the first 100 — 75%.

Spread across directions

Logarithmic scale. Dashed line — effective dimensionality (participation ratio) 64: this many independent directions are actually used instead of 3072.

Cumulative share of spread

To gather 90% of the dispersion, more than a hundred directions are already needed — the cloud is not completely flat. But it is not full-blooded either: 3072 dimensions degenerated into ~64 working.
Why this matters. Effective dimensionality is the width of the vocabulary that the engine actually uses. If it grew over time, one could speak of an expansion of the territory. Across the engine phases it stays in place: 47.0 → 46.0 → 49.7.
step 5 · step
What one step looks like

From one work to the next, the vector shifts: Δ = v(next) − v(current). To see this shift, it has to be projected somewhere, but not “for beauty”. The plane below is constructed honestly: horizontal — direction of the previous step, vertical — what remains of the direction “toward the center of the cloud” after subtracting the horizontal. Both coordinates are real scalar products in 3072 dimensions.

Each point is one real step of the corpus (700 total), color is epoch. Hover — you will see the plot to which the step led.

What the cloud shows

Points lie on the left: the next step, on average, goes against the previous one. The average angle between neighboring steps 119° — in free wandering it would be 90°.

Points lie from above: the step is systematically directed toward the center of the cloud, average cosine +0.63. This is a restoring force — a weight on a spring.

Together, this is exactly “trembling on a tether”: the step amplitude is large, yet it does not let it go anywhere, because each next one plays the previous one back.

Notice what is no: the cluster at bottom right. Without the restoring force and turnbacks, the cloud would be a symmetric circle around zero, and then the plots would diverge through space instead of crowding in place.
step 6 · reverse translation
From a point in space — back into text

Embedding irreversible: from 3072 numbers, it is impossible to reconstruct the text that produced them. The space can be read in only one way: take a point and see which real work is closest to it. Below is movement along drift axes, that very single coordinate along which the corpus moves over time: nine stops from the early pole to the late pole, with a real corpus work at each one. Move the mouse from left to right and read how the theme changes.

——
Between works — emptiness. If you take a straight line between two extreme works (gen 174 and gen 718) and walk along it, then at 13 stops out of 13 the nearest real work remains one of the endpoints: in the middle of the path the miss reaches 0.176, and the nearest work, apart from the two endpoints, is at 0.227. The space is populated with spots — the engine does not traverse the territory, it jumps across inhabited places.
This is where the junction happens. Meaning lives in geometry, because nearby points are read as nearby texts. This is exactly what gives the right to measure thematic drift by distance. But the engine itself does not make a move in this space: a point in the middle of emptiness is not text, and no formula can be written from it. A move is made where it is defined — in a discrete grid. This is the next screen.
step 7 · where the move is made by a formula
Grid: material × composition × palette

The engine has its own dictionary of visual decisions: 20 materials × 10 compositions × 8 palettes = 1600 cells. The work receives a cell, the genotype is written as m4_c8_p8, and the next work gets a neighboring one — and the engine records this transition literally as the operation palette:6->2. This is the “step by formula”: it is made not in semantic space, but here.

Material × Composition Grid · color = how many times the cell is occupied

How the move is calculated

Each axis is just a list. A move is a change of index on one or several axes: palette:6->2 means “take the sixth palette and replace the second.” Across the corpus this changed as follows: one axis — — times, two — —, all three are — —.

Visited 564 cells out of 1600. Where to go next is decided by MAP-Elites: it maintains an archive of the best works by niche and prefers rarely occupied ones.

The vocabulary grows. Indexes beyond the base lists are what the agent invented for itself: beyond the base, it grew 10 materials, 6 compositions and 4 palettes. Their texts are not in the dump, so in the examples such values are marked as “grown”.
step 8 · chain
Cell → medium → plot → prompt → image

Here is the complete chain of one real work: from the move in the grid to the finished image. Each block is real text from the generation trace; nothing has been reconstructed. Here you can see the place where the formal move turns into an intention: between the cell and the plot stands a language model that must link its assigned material, composition, and palette into a meaningful scene.

step 9 · perceptual space
The third space — and what all this means

The finished image is also vectorized, but by another model and according to another feature: CLIP turns image into 1024 numbers, where what is nearby is what is similar visually, not by meaning. In this space, the engine holds 64 niches (MAP-Elites) and ensures the works do not bunch together. In the dump, CLIP exists for 517 works.

Perceptual map · point = image, rings = niches

Projection of the CLIP space onto the first two principal components. Point color is time. Hollow circles are the centers of 64 MAP-Elites niches; occupied 64.

How much meaning reaches the image

Each point is a pair of works: horizontally, the distance of their plots, vertically — the distance of their images. There is a connection, but it is weak: ρ = +0.21.
0.21
distance between two images one and the same plot
0.29
distance between images different plots
0.56
drift over time in perceptual space (in semantic space — 0.74)
64/64
niches occupied — by images, the corpus covers the whole field
One plot — two images, and they diverge almost like strangers (0.21 versus 0.29). Text underdetermines the image: the lion's share of the visual decision is made after the meaning has already been formulated: in the medium, in the prompt, and inside the generative model itself.

This is computational creativity

Creativity here is not the moment when the model “invented.” It is a closed loop of three spaces. In grid move is defined formally, and it can be made and recorded. In language a move turns into an intention: medium, moment, plot — what was not in the grid and what cannot be inferred from indices. In perceptual space the result is measured as a picture and returned back into the archive of niches, determining where the system will go next. And the semantic space to which the entire rest of the report is devoted is our measuring instrument on top of the loop: in it, you can see where the theme actually moved.

Creativity in such a system is not a property of the step, but of the circuit: the ability to keep the search open without collapsing into either repetition or noise. This report measures how well the circuit copes. The answer on the full corpus is honest and bleak: the step span is huge (80% of the distance between random works), but each next step plays the previous one back (119°), and the cloud pulls toward the center (+0.63) — the search is moving, the territory is not growing. Meanwhile, the tether itself is drifting: the drift axis correlates with time at 0.74, and over 743 generations the theme has passed — eras. There is creativity here — but it is not expansion, it is the slow demolition of an inhabited area.

1 · Plot Map and Center-of-Gravity Trace

Each point is one plot in semantic space (UMAP by cosine). Gray — the entire corpus, bright — active in the selected window. Polyline with gradient — the center-of-mass trace over time: from blue (start) to yellow (finish), that is, where the theme drifted. The green circle is the center of the current window. Individual works jump almost randomly, so what you need to look at is precisely the trace, and even more reliably, the drift axis in the next section.

Map

Generation plot

2 · Drift axis — systematic displacement under noise

The single direction in 3072-dimensional space along which plots change monotonically over time (found as the covariance of the embedding with the generation number). Points are individual works; the line is a moving average over 15 generations. The scatter of points is huge, but the line moves confidently upward: . The poles of the axis are labeled with the words that correlate with it most strongly.
Early pole (bottom of axis)
Late pole (top of axis)

2b · What this axis means — three independent readings

A separate embedding coordinate means nothing: meaning in this space is recorded directions, not coordinate axes. The drift axis is exactly such a found orientation, and it can be read. Here are three methods that were computed independently of one another. First — project the words themselves onto the axis: 938 words are encoded by the same model, as the plots, and laid out along the axis by their own position in space — an answer from within, without referring to the corpus texts. Second — to see which tf-idf term frequencies grow together with the work coordinate. Third — simply read the names of the extreme works. If the three methods converge, the axis has been read correctly; divergences are also shown.
Axis ruler: ordinary words (not corpus jargon — a list of general concepts like "skin", "machine", "memory", "dust") are laid out by their projection onto this direction, in σ deviations. Corpus terms are intentionally absent here: they sit at the extreme edges and do not show the middle of the axis. Read the order from left to right — zero itself means nothing.

lower pole of the axis
1 · words projected onto the axis
2 · terms whose frequency grows toward this pole
3 · extreme works of the corpus

upper pole of the axis
1 · words projected onto the axis
2 · terms whose frequency grows toward this pole
3 · extreme works of the corpus

3 · Epochs and Real Works

The boundaries of epochs are found by binary segmentation on embeddings: a split is accepted only if the gain in explained variance exceeds the 95th percentile of the same gain on the time-shuffled corpus. Under the band are the works themselves; hover to enlarge.

3b · Plot tape — reading the trajectory step by step

The same drift, but in words. Each row is one generation: the plot title, the summary, the execution technique, and the size of the step from the previous work. Reading straight through, you can see how vector displacement turns into a change of theme: where the text crawls by synonyms, and where it truly breaks. Click on a row reveals the full summary and highlights the work on all charts. The “fractures” filter leaves only generations at epoch boundaries and the largest steps.

4 · Mathematics of Motion

Six checks answering one question: is this directed drift, random walk or trembling on a tether. All null hypotheses were tested by permuting time (500 repeats, fixed seed).

5 · Return Matrix

Cosine of each work with each. The diagonal is itself with itself. Bright spots along the diagonal = epochs (neighboring works are similar). Bright spots far from the diagonal = return: the agent returned to a theme it had already passed through. Hover to see the pair; click to go to the work.

All pairs

Hover over a cell.

Novelty: how unlike anything in the past the plot is

For each generation — 1 − maximum cosine with all previous works. Failures = the agent repeated what had already been traversed. Red markers are the nearest “double” in the past.

6 · Semantics: what exactly shifted

7 · Shift and Quality

If a large jump is research, it should pay off in the critic's score. We check this with Spearman rank correlation.

Step size → estimate of the next work

Novelty → estimate of this work

8 · Frame and Plot · Engine Phases

Series Thesis versus Work Plot

Step size by engine phases

What each phase is occupied with

9 · Method and Caveats