Some narrative and motif data in the Folklore Database is derived from Ю.Е. Березкин, Е.Н. Дувакин, Тематическая классификация и распределение фольклорно-мифологических мотивов: Аналитический каталог, referred to here as the Berezkin Analytical Catalogue. This source is made available under CC BY-NC-SA 4.0. Where Berezkin-derived material is used, the original source should also be cited.
The original Berezkin material is primarily in Russian. For inclusion in the Folklore Database, this material has been translated into English using DeepL and then manually checked and refined where errors are identified. DeepL was used because it allows HTML code to be preserved during translation, reducing both translation time and processing costs.
Although the Berezkin Catalogue contains a formatting structure, that structure is not always consistent enough to allow fully automated parsing. Automated extraction without correction can therefore result in corruption of narratives, traditions, regions, motif assignments, and associated metadata. To reduce this risk, the data is first structurally normalised before being parsed into the Folklore Database.
This structural normalisation is carried out using a locally run large language model, currently Qwen 3.5 35B, rather than an external public AI service. This approach is used to reduce the risk of transmitting third-party source material, translated catalogue content, or database-ready material to commercial AI platforms.
Once the structure has been normalised, the data is parsed using in-house Python tools and added to the Folklore Database only where it is deemed fit for purpose. During this process, translation, formatting, and parsing issues are identified and corrected before public release.
The narrative summaries in the Berezkin Catalogue are often highly condensed. This form is useful for identifying motifs and mythemes, but it can reduce readability for users such as storytellers, writers, students, and researchers seeking narrative flow. For this reason, many narratives are enhanced into a more readable form. These enhancements follow strict editorial rules: the original plot structure, motif content, mythemes, sequence of events, and culturally significant details must be retained, and no new narrative elements should be introduced.
Narrative enhancement is initially carried out using a locally run large language model, currently Qwen 3.5 35B, and is then manually checked and corrected where necessary before publication. The original translated data is retained separately and linked back to the relevant entry in the Berezkin Analytical Catalogue wherever possible.
Enhanced narratives should therefore be understood as curated and edited versions derived from the Berezkin source material, not as verbatim translations of the original Russian catalogue entries, although these translations are also available within the Folklore Database.