Garbler Devlog 02

« RYON EVERETT'S CYBERSPACE JUNKYARD

Dev log

The long dark tea time of the (audio-rendering on)sole.

Previous Devlog

Back to Vox Garbler Page

See more Devlogs

We (don't) have boundary detection at home

If you've been through an MRI sans plugs or music: you know what the beast of two-fax-machined-backs is supposed to sound like. My expectations were at least that low for my first boot: Some rhythmic cadence representing a larger force at play. Even if, while strapped in and getting it broadcast magnetic resonance screeches to the dome you couldn't quite figure out, and it was scary, you knew that pushing through held some value. An image, some clarity, the clanking was a vehicle for success.

I wish were my expectation were lower.

This was the inbred homunculus-robot hybrid crying out "existence is pain" in 8 to 32 bits through a manhole cover. Totally unintelligble. After about 100 renders and 20 different phrases all amounting to varying degrees of "igtomp -uhtreasfur paimuhdipa ooh(n)g". I realized this was not a smooth sail.

WHY SPIRITBOXING SUCKS

If you, like me, were relatively unexposed to the incredible science behind transcription models, you'd be forgiven, unlike me, to not know that they don't "just work." You see, I had a hunch it might get more complicated than I anticipated: I'd been exposed in some of my studies to linguistic softwares for studying speech. Formants, phonemes, and boundaries are just things that speech has, jargon standing between me and my glorious gamey talk-box.

It turns out, that the reason it took so long to have half decent audio-transcription at half-decent speed, is because it wasn't very easy to do it well in the first place. I naively thought that if I just ran the most sophisticated ASR (audio speech recognition) model at my tiny 800 000 word bank of audio, it would mostly accurately find exactly the timestamps for boundaries.

Crash course using ASR transcription AI models for spiritboxes

101-word boundaries are exceptionally difficult to extract cleanly. Most of the bright-minds agree that phone boundaries are most accurate in the middle, not at the edge of the word. Spiritbox nil, sanity -1

101-so nice we use it twice: audio quality (aka not shitty compressed audio from 30 year old games), speaker consistency (aka not voice lines that change speaker in a domain pretty much every other sample) and language content consistency (aka not a mix of multiple languages, spoken dialogue, grunts, sfx and music) matter a lot. Most models are not trained on voice over lines from games, so it often sounds like not language to them. Spiritbox: 0, uh-oh count: ++++++++

101-you-should-have-researched-this-dummy: the local open-source models and cloud models do not optimize for ms accuracy in their transcriptions, because they're not designed to feed sample-based spiritboxes (who could have guessed). Spiritbox: broken, will-to-continue: injured.

Screech by screech

I knew my body of audio had many different flavors, yum, but that presented a challenge: which tool was closest to getting good timestamps for boundaries, and which levers could I pull to improve my render quality.

The great listening bake-off

I knew I had to do several things if I wanted to turn the audio into spiritbox-friendly audio.

1- I needed to isolate english audio vo so that our renderer didn't receive background music, sfx, other languages, that were transcribed by our first ASR model as having words

2- I needed to test more boundary writing transcription models to see if I just picked a dud

3- I needed to know if it were even possible to slice near word boundaries accurately, or if I had to rebuild the whole concept around phonemes (more on that later later)

So with the problems in mind I hatched a scheme: Triage audio at the source based on file-structure and pattern to pre-index likely category audio belongs to: english, japanese, sfx, bgm, misc, and cutscenes (which need to be handled differently).

I decided on a blended approach of taxonomy semantics (where the audio is inside a directory, whether that directory seems to contain a certain type based on the structure of the file system and the game it's a part of. That gave me a first pass canvas label saved to a lil database.

Then I ran triage models on local inference. Their jobs were simple: listen to an audio sample. Label if you think it's spoken language. Label what language you think it is. Label if it's not one of the two languages I had in abundance in the data (English and Japanese).

After the several day triage I prepared a testing and filtering applet to evaluate the taxonomy labels and triage agents' accuracy. It used evenly distributed sampling across each directory with audio, So of my 800 000 words spread over 136000+ audio files. Even my sample pool was quite large. The filtering app had 5 qualities to label.

The process fairly straightforward. Label and score each audio clip, comprate with automatic and triage agents. Find disagreement patterns. Tune the algorithm per domain. Rinse repeat. Eventually taxonomy and triage models would accurately predict the type of audio based on the previous input signals and I could finally go finally begin to have a labelling system that would signal what TYPE of audio was in a clip, accurately.

Now I needed to make the boundaries and quality gates less offensive.

The hole was deep, so I needed to mine

And on the I-lost-count-eth day, there was mining. But before that, there was a lot of A-B...N benchmark comparisons. I had to manually slice a pool of statistically significant samples by hand, which was tedious, but gave me a "golden" reference I could compare the plethora of ASR models with.

So after bundles of research testing cloud models, and then several weeks of my GPU raising the ambient temperature comparing the top 15 ASR local models for accuracy. It became clear to me that close-enough was not going to cut it. Off the shelf was store-brand in a not good way. I used aggregates and statistical averages of the models some models were good with As and not Es, I tried to isolate the accurate domains from each model's each run, and then weight them against my reference to split the difference.

After many experiments, the consistency just wasn't good enough using them stock/multi-models.

So, depesrate to not do what finally transpired, I tried another angle, optimizing the renderer: I played with renderer levers: leave some time before, add more time after, speed it up, slow it down, render contiguous stretches from a single sample before switching. They made modest improvements, but none could overlook the lack of accuracy for word boundaries.

Enter Machine Learning Science: Audio Edition.

Adding more AI wasn't going to cut it. I needed better AI: I needed fine-tuned/domain adapted machinery to get through the next hurdle, if I didn't want to manually slice 800 000+ words over the span of several years of not-working time.

If the words Viterbi, Lattice, Markov, Gaussian Mixture, Transducers, Conformers, are new to you, most were to me in the audio context. There are many nuanced mathematical ways to capture and express words and sounds, which makes the audio models genuine marvels of modern technology. Techniques to optimize detection of shifts in spectrographic information, comparing sequences, genuinely such a fascinating field. I needed to learn enough to be able to build something workable for my audio bank.

Enter reinforcement learning/fine-tuning and Montreal Forced Aligner(MFA).

In tests with arbitrary values, I could pair open models with MFA and tune the MFA parameters and it would move the boundaries around. Now that I knew the machinery would steer the models, I had another problem. The reinforcement learning loop and tuning required golden data samples, but not just a hundred and change like I did for my initial bot vs me challenge. I needed hundreds, to thousands, of sliced audio samples.

So I build a game make it less miserable to label everything and made it online so I could coerce my friends to help.

Into the depths of Garbler Mines