Top
Best
New

Posted by epestr 6 hours ago

How I reverse engineered a commercial spatial audio effect(epestr.com)
78 points | 27 comments
fxtentacle 2 hours ago|
It's been 12 years, but I once wrote a blog post on how my commercial spatial audio effect works: https://hajo.me/blog/2014/12/28/how-surround-sound-for-headp...

BTW, the resulting plugin was used during mastering of George Lucas' Red Tails, so I guess sound quality was fine ;)

My own Mac app has long since been discontinued (because Apple locked me out of the KExt permission), but I sold the tech and it lives on.

I keep wondering if I should build hardware headphones with the spatial audio inside, so it looks like Dolby Atmos to the TV but then does virtual surround sound on your headphones, but hardware has low-ish margins so I never found anyone willing to fund it, even though I already have a working prototype. Oh and the audio engineer who invented it has since been Grammy-nominated twice. And, well, Hollywood likes it. But still a tough sell towards investors.

andai 5 hours ago||
Interesting article. I don't know much about sound processing but it reminded me of the Fruity Convolver in FL Studio, which can somehow apply the characteristics of one sound (an impulse?) to your audio stream. It has interesting effects built in, reverbs recorded in different unusual spaces, as well as guitar amps.

I would have liked to compare the audio clip from the beginning of the article, at the end too! Instead I had to record my own voice. And there doesn't seem to be a way to toggle the effect on and off? So it's a bit hard to get a sense of what it actually does.

This also reminds me of this cool mp3 from the 2000s — maybe it's on YouTube these days — called virtual barbershop, which was recorded with binaural microphones.

Edit: here she is! https://youtube.com/watch?v=IUDTlvagjJA

epestr 5 hours ago||
Yup, that would be an impulse response, and if the room acts as a linear (and time-invariant) filter, then you can easily figure all that out, and the write-up is basically about testing if it's so and how good the reconstruction is.

And there is a way to toggle it: the icon which is orange can be tapped on and off, and you can adjust the "bass" which is built in with the ring around it. Also, it's likely to have no effect on voice, more complex arrangements will have an audible difference.

There is similarly an example on Wikipedia too, which got me interested: https://en.wikipedia.org/wiki/File:BinauralPaper.ogg.

_joel 5 hours ago||
Yes, it's called impulse response, used to capture convolutional reverbs, to model "real" sound stages like a large hall or what have you
clzerv 5 hours ago||
While I appreciate the effort and talent to pull this off, I find the results to be awful. I call this kind of overblown sound processing "Bose barf" since pretty much all Bose audio sounds like this, and it just grates on my nerves. Now TVs, Bluetooth speakers, headphones, etc. all seem to perform this kind of aggressive audio processing to overcome the limits of super cheap tiny drivers. It's getting hard to find devices which have a more neutral sound profile. All this audio processing makes music sound like it's coming out of a garden hose duct taped to a 1950's grade school hall speaker. I can't understand the appeal.
epestr 5 hours ago||
There's a subwoofer slider (around the icon) if the bass feels too heavy, and some songs require it to be toned down if it's already processed to be bass-boosted.

My Jabra headphones don't seem to apply any processing by default, and I prefer how they sound with the processing. If your headphones already boost bass or apply another effect, stacking the processing can be grating (which I've experienced when I record a video and play that through the daemon running).

I guess it's a matter of taste. I just like music sounding as though it's coming from a little farther away, and I suppose the bass-boost might be what you're finding unpleasant.

ryankrage77 4 hours ago|||
Seems to depend on the source audio too. The demo at the top of the page sounded good, but in the demo at the end I tried some rock music that already has a pretty wide stereo image, and it made it sound very 'squished' and like it was coming from inside my head, on top of lots of weird distortion/phase issues.
epestr 3 hours ago||
That sounds intriguing. I've only noticed it making sounds feel wider, or at most do nothing rather than squishing an already wide sounding track. What track were you trying, I'll see what's going on.
ryankrage77 2 hours ago||
It was 'Square Hammer' by Ghost.
tancop 4 hours ago|||
Yeah the default bass boost is too much. I set it to point straight left and its sounds way better.
epestr 4 hours ago||
I skipped this detail in the blog, but the spatial filter in the original processor already had bass boosting within the effect which I carried over and interpolate (non-linearly of course) now.

I've made the default 20% in the page for now.

diydsp 5 hours ago|||
Same here wrt to Bose. I respect their tech and science but then they add a bogus aesthetic layer on top.

I think it is meant to make jazz sound good. The eq-d bass and attenuated mids make a standup bass more apparent and the piano and vocals more present.

But it made me start to wonder if I started hating my favorite band: I was packing up to move this weekend and stumbled on my CD collection. Many Stereolab CDs. I played them in my basic acoustic waveguide radio. A slick bookshelf I inherited and rarely use.

45 minutes later I was telling myself "god this music sucks...just agonizing repetition." Turns out yes the motorik bass is repetitive...but the other key point of the band is the gentle singing and intricate delicate synth and organ melodies...all foreclosed on by some evil twin combo of an audio eng and marketer overindexing in mid low bass.

jeffreygoesto 3 hours ago|||
That is why I refurbished my Harman-Kardon amp from 1991 and listen CDs and LPs through it increasingly these days. It drives either some Revox speakers I be also bought at that time or a HifiMan Edition XS. That thing is outright amazing. The pinnacle of mixing was in the 80s as far as my collection is concerned. The 70s were much experimenting and the 90s already too compressed.
fwlr 5 hours ago|||
You’ll want to get something from Etymotic, they target the “precise and neutral” profile you’re after.

OP wants the `earwax` audio filter in ffmpeg.

mrob 4 hours ago|||
Any headphone/earphone can be EQed to "precise and neutral" if the distortion is low enough. IEMs like the Etymotic IEMs often do well here, but if you dislike things inserted in your ear canal you might prefer planar magnetic headphones. These typically have lower distortion than dynamic drivers, especially in the bass. I have a set of HifiMan HE400se headphones, which used to be exceptionally good value before the price went up a lot recently.

To get the best possible results from equalization you need to do it by ear by comparison to a calibrated reference loudspeaker, but you can still get good results using other people's measurements. And there's no consensus what the best target curve is. Studio control rooms aren't anechoic chambers so "flat" won't necessarily give you the intended sound. You might find AutoEq useful:

https://autoeq.app/

I generally prefer no spatial processing, except in the case of 60s-style hard-panned stereo, which is intended for loudspeakers and does sound a lot better with the "earwax" filter or similar. "Precise and neutral" is a lot more difficult with loudspeakers, both because it's hard to keep distortion low when you're moving a lot of air, and because you have the off-axis response and room acoustics complicating things.

Aurornis 2 hours ago||
> Any headphone/earphone can be EQed to "precise and neutral" if the distortion is low enough.

EQ can help a lot, but there are limits. The response curves you see are usually very smoothed and they’re only from one sample. The peaks and valleys and resonances can be narrow and can differ from unit to unit or even more from one manufacturing run to the next.

The response is also a function of the interface to your ear and leakage to the outside world. Some headphones are more sensitive than others.

I have a few high end headphones and I’ve tried a lot of the published EQs for all of them. In theory they should sound like they have the same frequency response but they still sound very different. Closer together when EQed, but you can tell they’re not the same.

I am an EQ proponent, but doing EQ on headphones doesn’t level the playing field so much that it only comes down to distortion.

The best case for EQ is in a sound controlled room for speakers when you can sit in the sweet spot and reflections and room effects are not a problem. Set up multiple sets of decent (distortion free) speakers in that environment with a moderate EQ and my listening skills aren’t good enough to identify differences most of the time.

mrob 1 hour ago||
True, and that's why you have to EQ by ear for best results. But even that won't be perfect, because the frequency response changes depending on exactly how the headphones are seated on your head, or exactly how the IEMs are inserted. And if there's foam in the construction, it will gradually compress over time, also changing the frequency response. It's still likely to be better than what you can get from loudspeakers in a typical room.
epestr 5 hours ago|||
Thanks, I hadn't come across earwax before. I listened to the examples here: https://hhsprings.bitbucket.io/docs/programming/examples/ffm...

It sounds very close to what I was going for, and feels much softer to hear. I wasn't able to find anything similar earlier with any search engine or an LLM.

maybewhenthesun 3 hours ago||
100% agreed.

This fashion of cranking the base and the high end and removing the middle (simulating a subwoofer with some tiny speakers, basically) can't die fast enough.

I mean, fair play for the writer. If he wants it to sound that way he's of course free to do so. But this style is used a lot in concerts and open air amplification and I hate it.

My theory is that this style is popular with people playing their music too loud. If your ear feels 'blocked' after a while, your music is too loud.

One thing with this style is that, because it leaves the middle (speech) frequencies relatively empty, it feels less blocking. And the high base makes that it doesn't really sound tinny or harsh. But they high base cane easily damage your hearing still.

epestr 17 minutes ago|||
OSHA uses A-weighted sound pressure to estimate hearing risk, which roughly follows the ear's sensitivity and gives low frequencies much less weight, much like Fletcher Munson curves. Check section 10 here: https://www.osha.gov/otm/section-3-health-hazards/chapter-5. So bass having more volume/power doesn't by itself mean it's equally damaging to the ears.

I also wonder about the opposite situation: someone turning the whole volume up because the bass doesn't sound loud enough, raising the other parts with it. Having the bass adjustable could help there. Some songs already have plenty of bass, so it makes sense to turn it down for those, and perhaps have a dynamic limiter when it gets excessive.

By "blocked" I meant the impression of music sounding too close to my ears, rather than my hearing feeling blocked after listening. Of course, feeling softer doesn't completely establish that it's safe either.

TheOtherHobbes 1 hour ago|||
It's not a new fashion. In the 70s/80s hifi often had a "loudness" button which boosted highs and lows, although not as much as they're boosted by EQ settings today.

The more recent variant is "smiley curve" EQ - high at the extremes, drooping in the middle.

EQ can mess with transients, especially in the bass. You get more volume, but at the expense of some time smearing.

coldblues 3 hours ago||
In hobbyist audiophile communities, there is a focus on downmixing Dolby Atmos music, which is 5.1 surround and has a much higher dynamic range. As an example: https://github.com/peqdb/macos

If you have Apple Music and you're on Windows, in the settings there is an option to downmix to stereo. Give it a try on any of the music here: https://helloatmos.app/library/

There are also people experimenting with spatial audio using HRTF, but I don't think it's very good for music.

haddr 2 hours ago||
I wonder if author could follow something similar that is used to capture guitar amps to NAM models (Neural Amp Modeller), https://www.tone3000.com/capture
schleck8 3 hours ago||
I'm not convinced, the Nothing Ear spatial audio effect (which is pure postprocessing as far as I understand, since it's independent of codecs and containers) sounds a lot more realistic, and I doubt Nothing has the best algorithm in the field
epestr 28 minutes ago||
Further, what's technically the most correct, like SADIE or MIT are the least impressive sounding ones. I suppose the question then becomes which sounds the "best" and remain subjective.
epestr 3 hours ago||
I'd be interested in comparing it with the Nothing Ear effect. If it's linear (and time invariant), it should only require playing a test sample through it, recording what happens, and using an automated script to extract the responses.

As for the effect, I was trying to get a particular effect I'd liked using earlier, rather than the most realistic spatial audio algorithm. The earwax effect mentioned elsewhere in the thread also sounds quite close to what I wanted. Posting this also worked as discovery, getting me some good suggestions for other effects to reverse engineer.

jonathanstrange 1 hour ago||
Sounds awful but I appreciate the technical aspects of the post.
aarunsoman 4 hours ago||
[flagged]
ofabioroma 2 hours ago|
[flagged]