Latvju dainas · 46,587 songs · a corpus study

The Cabinet, Counted

Krišjānis Barons spent most of his life sorting Latvian folksongs into a cabinet of drawers. Four measurements, a century after he finished, ask what his system knew.

Over some four decades, one man read a quarter of a million Latvian folksong texts, one at a time, each on a slip of paper three centimetres by eleven, and decided where every one of them belonged. The instrument he built to hold them — a cabinet made in Moscow in 1880, seventy-three drawers, 268,815 slips — is the Dainu skapis, on the UNESCO Memory of the World register since 2001. The edition that came out of it, Latvju dainas, ran to six volumes between 1894 and 1915 and holds 217,996 songs.

Barons was clear that the sorting, not the collecting, was his contribution. It is not a book that happens to be organised. The organisation is the work.

Latvian folkloristics has read the dainas overwhelmingly through close reading: quote a song, read outward from it. That method has produced nearly everything worth knowing about them. What it cannot do is look at forty thousand songs at once.

This page does that, on the slice of the collection that is machine-readable. Four questions, in order of how much each leans on the one before it: where the reusable language sits, what a song does at its start, where variants of the same song come apart, and — the one that turned out to be about Barons rather than about the songs — whether the words agree with how he filed them.

What is being counted

The corpus is published by the Institute of Mathematics and Computer Science at the University of Latvia at korpuss.lv: 46,587 song texts, 921,508 words, recorded 1770–1903, lemmatised, and — the thing that makes any of this possible — each song still carrying Barons’ own hierarchical classification, with his Latvian descriptions intact: 2.4.4.2.1. Meita tura godu, sargā vaiņagu, netura godu, zaudē vaiņagu.

It is one volume of the six, not the cabinet. Specifically it is Volume 2, dated 1903 — the youth songs. Every classified song carries a code beginning 2.4: adornment, the wreath, the ring, learning to work, the dowry, reputation, girls’ and boys’ songs, night visits, old maids and bachelors. Childhood, weddings, married life, death and the mythological songs are in the other five volumes. Everything below is true of the youth songs and is not claimed of the rest.

One more gap, invisible in the data: the printed volume holds 47,781 texts and this corpus has 46,576, because the Latgalian-language songs were excluded when it was built. So it is not quite complete even for its own volume, and the exclusion removes a dialect region — worth knowing before reading anything here as a statement about Latvian folksong at large.

46,587songs
47Barons sections
270,973distinct 4-word runs
1770–1903recorded

AWhat travels

The formula inventory

Oral song is built from reusable pieces — that has been the working assumption since Parry and Lord went looking for Homer in Yugoslavia. The question a corpus can settle is which pieces, and how far they get.

“Reusable” needs more care here than it first appears to, and getting it wrong is how this section was originally published with the answer backwards.

The obvious measure says almost nothing travels. Of 270,973 four-word runs, 98.3% never leave a single theme. Restricting to runs common enough to have had the chance — the 6,060 appearing in ten or more songs — still leaves 81.8% of them staying put. That looks conclusive.

It is an artefact, and the corpus's own structure is what produces it. This archive recorded the same song over and over from different singers, and Barons filed variants of one song into one section by definition. So a run appearing in fifty songs may be appearing in fifty recordings of a single song, and its staying in one theme says nothing whatever about formulae. Frequency does not screen this out: a run in “ten or more songs” is usually still one song written down ten times.

The control that works asks whether the songs sharing a run are actually different songs — measured as how much of their vocabulary they share overall. Split that way, the picture reverses:

Songs sharing the run are…RunsStay in one theme
near-duplicates of each other (a variant family)4,49388.9%
moderately similar1,27368.0%
genuinely different songs29432.7%

Three quarters of the “frequent” runs are variant clusters, and those are the ones that stay put — only 11% of them cross a theme, which is close to a restatement of how Barons filed. Strip them out and just 294 four-word runs are genuinely shared between different songs. Of those, two thirds cross two or more themes.

The formulaic core of this corpus is small, and it is mobile. What looked like theme-locality was duplication wearing its clothes.

This matters beyond the one number. It means most repetition in the archive is not a singer reaching for a shared phrase — it is a collector writing the same song down again. Which is worth knowing before treating any frequency count here as a fact about how people sang.

And the ones that do get out

ThemesSongsSequence 
1099kad es biju jaunawhen I was young
1096es biju jauna meitaI was a young girl
1023ai dieviņi ai dieviņioh God, oh God
943vai dieviņi ko darīšuoh God, what shall I do
817kad es iešu tautiņāswhen I go to the suitors’ house
765es mātei viena meitaI am my mother’s only daughter
747visi saka visi sakaeveryone says, everyone says
712visi ļaudis tā sacījaall the people said so
649es bagāta mātes meitaI am a rich mother’s daughter
645drebi drebi apšu lapatremble, tremble, aspen leaf
632vai māmiņa vai māmiņaoh mother, oh mother

Read the list and the shape of the system is visible without any further statistics. What travels is frames, not images: self-introductions, addresses, laments, the “everyone says” that sets up a report of gossip. This is the machinery for getting a line started and pointed at a subject.

The images stay home. The wreath, the oak, the dowry chest, the mill — the objects the songs are actually about — do not appear on this list at all. What crosses Barons’ boundaries is grammar under pressure of metre. What stays inside them is meaning.

BHow a song starts

The opening formula system

A companion analysis on the dissertation site found that the first tenth of a daina is 2.3× more formulaic than everything after it. That says the opening is a distinct functional slot rather than just the place a song happens to begin. So: what is in it?

20,619distinct openings
69.2%of songs share theirs
27.2%covered by the top 500

Seven songs in ten start with a line some other song also starts with, and the 500 commonest openings — 2.4% of the inventory — account for better than a quarter of the whole corpus. There is a long tail behind that, but the head of it is small and heavily used: enough openings that a singer always has one ready, few enough that every one of them is familiar to the room.

And the slot is genuinely reserved. Of the 1,665 openings established enough to be used by five or more songs, only 11.8% ever appear anywhere but at the start of a song. These are not general-purpose lines that happen to turn up first. They are lines for starting with.

SongsOpening 
177dod māmiņa kam dodamagive me away, mother, to whoever you will
163ābelīte dievu lūdzathe little apple tree prayed to God
160pieci brāļi viena māsafive brothers, one sister
121adu cimdus adu zeķesI knit mittens, I knit stockings
98kad es biju jaunawhen I was young
88kalniņā stāvēdama lejā laidustanding on the hill, I let it down to the valley
86jauni puiši mutes devathe young lads gave kisses
80ganos gāju kreklu šuvuI went herding, I sewed a shirt
76es meitiņa kā puķīteI am a girl like a little flower
73caur ābeļu birzi jājuI rode through the apple grove
70puķe puķe roze rozeflower, flower, rose, rose
67protu protu redzu redzuI understand, I understand, I see, I see

Notice what these are. Almost every one establishes a speaker and a situation in a single line — who I am, what I am doing, what I am asking for. Two of them (puķe puķe roze roze, protu protu redzu redzu) do nothing but establish rhythm and a voice, saying essentially nothing at all. Those are warm-up lines in the most literal sense.

CWhere a song lets go

Variant drift

The archive records the same song many times over. That is usually treated as a nuisance — noise to be deduplicated before the real analysis starts. It is also a natural experiment: every family of variants is a record of what singers kept and what they changed.

Group every song by its first verse line, keep the families of five or more, and ask how far the family agrees, line by line. 1,702 opening lines qualify, covering 22,204 songs — nearly half the corpus.

0% 50% 100% line 1 line 2 line 3 line 4 100% 51.1% 46.2% 43.9%
Agreement on the commonest line within a variant family, by position. Line 1 is 100% by construction — it is what defines the family.

The collapse is immediate, not gradual. By line two the family has already split in half, and after that it barely erodes further — lines three and four are only a little worse than line two. Whatever happens, happens at the seam between the first line and the second.

Songs2nd linesShared opening
17566dod māmiņa kam dodama
17031ābelīte dievu lūdza
16046pieci brāļi viena māsa
12120adu cimdus adu zeķes
11316kalniņā stāvēdama
11123maza maza meitenīte
9544kad es biju jauna meita
9429nav neviena ozoliņa

A daina variant family is not a song with small changes. It is a fixed opening with an open continuation.

Which is the same finding as the front-loading of formula, arrived at from the opposite direction. The opening is a shared, public, reusable object — a hundred and seventy-five singers reached for dod māmiņa kam dodama and then went sixty-six different ways. Everything after the first line is where somebody is actually singing.

DWhat Barons knew

The taxonomy as an object of study

Barons had no theory of classification and no committee. He had slips of paper, a cabinet, and four decades. The obvious modern question is whether the categories he arrived at are real — whether they correspond to anything in the songs, or whether they are one man’s idiosyncratic filing.

The test: train a plain naive-Bayes classifier on the words of 80% of the material and see how often it recovers his filing decision on the rest. Forty-seven sections to choose between.

One thing has to be got right or the number is worthless. The corpus averages 2.5 recordings per song, so splitting at random puts variants of the same song on both sides and the classifier gets to recognise rather than generalise. The songs are therefore grouped into 18,215 variant families by shared first line, and whole families go to one side or the other. That correction is worth about eight points: the first version of this page reported 87.0% from a random split, which was inflated.

79.1%recovered from words alone
19.8%baseline (largest section)
47possible answers

It is still a high number, and that is the first thing worth saying. A taxonomy built by hand in the decades around 1900, by one man reading songs one at a time, is largely recoverable from vocabulary alone a century later — four times the baseline across forty-seven classes. Barons was consistent to a degree that would be unremarkable only if he had been following a written rule, and he was not.

The other 21%

The interest is in the failures, because the classifier does not fail randomly. It fails in exactly two ways.

It collapses sections that share a situation and differ only in who is speaking. Its commonest mistakes are girls’ songs, boys’ songs, and other people’s advice about courtship — three sections using the same words about the same events. What separates them is the voice, and voice is not in the vocabulary. Barons separated them anyway. He was tracking who sings and to what end — a judgement about the songs that the words cannot carry and that no clustering algorithm would ever have made.

And it fails on his finest distinctions. He divided the weaving songs into thirteen sub-sections by which garment is being made: shirts, mittens, stockings, shawls, skirts, sheets, belts, the bleaching, the fulling. Those are among the worst-classified sections in the corpus, because the songs are nearly identical apart from one noun.

RecoveredSection
94.6%2.4.7.1 Meitu dziesmas girls’ songs
94.1%2.4.3.2 Vaiņags the wreath
93.8%2.4.5.9 Maltuve the grinding room
92.9%2.4.6.1 Ļaužu valodas… gossip and reputation
 …
60.3%2.4.2.2 Seja, vaiga sārtums, acis… face, eyes, beauty
59.3%2.4.5.8.8 Auž, raksta, pušķo villanes weaving shawls
54.8%2.4.5.8.12 Audeklu balināšana bleaching cloth
48.5%2.4.6.2 Ļauni ļaudis, skauģi, burvji… ill-wishers, envy, sorcerers

That is not a flaw in his scheme. It is a record of how finely a nineteenth-century Latvian farm distinguished work that a lexical model — and most modern readers — would see as one activity. He kept a distinction the language itself barely marks, because the people he collected from kept it.

The cabinet is not a container for the songs. It is an argument about them, and most of it survives contact with the evidence.

Method, and a note on manners

Everything here runs on a single local snapshot of the corpus, harvested once and analysed offline. korpuss.lv is a public academic service run by a research institute, not a commercial API; firing a query per question at it would be poor manners, and analysing a fixed snapshot has the side benefit that every number above is reproducible. The code is a few hundred lines of Python with no dependencies.

Three things that had to be got right

Editorial parentheses are stripped. The printed edition records variant readings inline — Kam, meitiņa ( māsiņa), skaista augi. Those are the editor’s, not the singer’s, and leaving them in inflates apparent formulaicity exactly where an editor happened to notice variation.

Section-crossing, not song-crossing. In a variant-rich corpus, “shared between songs” and “written down twice” are the same measurement. Requiring a sequence to cross two themes is the difference between 27.2% of four-word runs looking formulaic and 1.7% of them actually being so.

Unpunctuated line breaks are recovered. The edition does not always set a comma at the end of a verse line, so splitting on punctuation alone silently fuses pairs of lines. The tell is a population of fifteen- and sixteen-syllable “lines”, which is twice a daina line and therefore impossible. Uncorrected, it distorts every count of song length and every variant family.

What this is not

It is one volume of Barons out of six, and roughly a fifth of the songs he collected. It is large enough for distributional claims and far too small to prove an absence: “not attested here” never means “does not exist.” Nothing above is a claim about Latvian folksong in general, and the mobility result in particular — a formulaic core of only ~294 runs — would almost certainly look different across the whole cabinet, where there are five more volumes for a formula to travel into. That is the next thing to find out.

What was corrected, and why it is left visible

Two results on this page were published wrong first, and both failed the same way: the corpus is variant-rich, and neither analysis controlled for it. Section A read duplication as theme-locality and reported the opposite of what the evidence shows. Section D reported 87.0% where the honest figure is 79.1%, because variants of one song sat on both sides of a random train/test split.

Both are corrected above, and the wrong numbers are left in view rather than quietly swapped out. The trap is a real property of folk-song archives — the same song recorded over and over is the normal condition of this material, not an edge case — and the next person to count anything here will meet it too.