After figuring out what a “linear phase speaker” means and checking some professional studio monitors in Part I, we can apply these ideas to the tuning process of my LXdesktop.
Before throwing in the power of linear-phase DSP processing and paying the latency costs, let’s first check what the original LXmini implementation looks like in terms of phase rotation and group delay. For that, we first build a simple idealized approximation of it and compare it with the actual measurement.
There are three major factors that define the phase rotation in the original LXmini:
All these components are minimum-phase filters. Modeling the first two is rather trivial. Modeling the low-pass part cost me some time. I figured out that besides the phase rotation from the low-pass filter itself, the other contributing factor is the brick-wall filter employed by the DAC/ADC that are used for measurement.
I discovered this because a minimum-phase low-pass filter of the same magnitude shape did not exhibit the same phase behavior, so there had to be an all-pass component. Measuring the electrical chain alone, at higher sampling rates, confirmed that.
Thus, the LXmini model I’ve built here is not purely theoretical: I had to fit the properties of the low-pass filter to my measurement. However, the nice part is that even the impulse response of this model looks very close to the actual IR at the macro level:
Note that since in the LXmini the polarity of the full-range driver is inverted, the pulse actually starts on the negative side. However, since acoustical software typically looks at the largest positive peak, to avoid confusing it I have changed the polarity of the entire pulse—this does not affect the magnitude response, and allows Acourate to find the peak and reference the phase to it correctly. This is the graph of phase shifts of the actual speaker and its model:
So the total phase rotation (if we unwrap it) is approximately 520° of which 180° comes from the high-pass, another 180° from the crossover, and the rest is the low-pass filter plus the measurement artifact from the anti-aliasing filter of the audio interface.
The resulting group delay predicted by the model:
(Note that Acourate places the start of the impulse at sample position 6,000; thus, the baseline group delay is approximately 125 ms.)
We can see that the group delay starts rising below 2 kHz and reaches approximately 1.6 ms at 100 Hz. This is actually not that bad (and definitely better than the Genelec 8331A in the “low latency” mode!) because presumably the group delay stays within inaudible boundaries, according to the research considered in Part I of the post.
We can conclude that the original LXmini is definitely well designed from the psychoacoustic perspective. However, it is not a “linear phase speaker,” thus some improvement may still be possible.
LXdesktop differs from LXmini in two aspects: the pipe of the woofer is shorter, and I use a sealed subwoofer. Thus, it is a three-way system with two crossovers to define. Let’s start with the lower one.
Choosing the crossover frequency between the subwoofer and the woofer has turned out to be a challenging task. We don’t want to overload the woofer by forcing it to go too low: due to the smaller enclosure volume it can’t go down to 45 Hz like the original LXmini. Since I’m tuning this speaker setup for my room, the choice of the crossover is largely dictated by the interaction of the woofer and the subwoofer with the room modes. Unlike the Dutch & Dutch speaker, my subwoofer is not co-located with the woofer, thus they couple to the room modes differently—in fact, significantly differently. This dictated the choice of the crossover: find the point on the frequency scale where their phase behavior is more or less close, and narrow down the crossover overlap region to that frequency as much as possible.
However, since speaker drivers are naturally minimum-phase systems, any deviations in the magnitude response affect the phase. Thus, the drivers must be linearized before doing the comparisons. Another important thing is that measurements of both drivers must use the same time reference: preferably a high-frequency driver for producing a narrow synchronizing peak.
We can see that the phases of both drivers (red: subwoofer, green: woofer) have similar slopes in the region around 60 Hz, so I ended up choosing 60 Hz as the crossover frequency. Then I started looking for a brickwall-like crossover to minimize the overlap region. With linear-phase crossovers even steep slopes by definition don’t pose any phase alignment issues; however, they have a potential issue of pre-ringing. The pre-ringing definitely must lie beneath psychoacoustic pre-masking thresholds. I considered several types of steep crossovers: LR8, 2nd-order Neville-Thiele, and 2nd-order polynomial UB (Brüggemann). All of them offered negligible pre-ringing, and UB are the steepest ones, so I chose that type. Here is how its pre-ringing looks:
The pre-ringing above -60 dB from the peak lasts for approximately 5 ms.
Note that under ideal conditions: when the speaker drivers actually behave like the crossovers, and with the listening point at the tuning position, the pre- and post-ringing parts of the driver IRs cancel each other completely. However, as we will see this is hardly achievable in reality, thus the pre-ringing of the crossover must be kept to the minimum.
For this problem, we have two constraints. First, since the woofer driver is oriented upward and is located approximately at ear level, we should limit its range to the region of its omnidirectional radiation pattern. According to the technical documentation on the SEAS L16RN-SL driver, it is omnidirectional up to 500 Hz. However, this is the area where the second constraint is not met—the full-range driver, used without a baffle, simply can’t go that low without risking over-excursion. Thus we need to pick a trade-off between operating the woofer driver mostly in its omnidirectional range while avoiding overloading the full-range driver. LXmini uses 700 Hz as the crossover point, and I decided to keep it.
Side note: previously I was trying to find a full-range driver that would exhibit the lowest distortion in baffleless operation. However, maybe a better criterion would be to pick the one which allows going as low as possible without going over its excursion limits.
Although I kept the crossover frequency, I also decided to narrow the crossover overlap region: the LR2 used by LXmini is too broad. I checked steeper crossovers first, and unfortunately they all had significant pre-ringing, which is a bigger concern here than in the bass region because human hearing is much more sensitive in the midrange. So I reduced the crossover order and ended up using 1st-order Neville-Thiele:
(Note that the time scale on this graph is different from the previous one.)
The pre-ringing above -60 dB from the peak lasts for approximately 1.2 ms. While temporal backward masking (pre-masking) is generally weak, it is highly effective within a brief 1–2 ms window immediately before a loud transient, so this pre-ringing should stay hidden.
With both crossovers defined, it’s time to take our mixed-phase ideal speaker model from Part I and split it into the bands for each driver. Due to the minimum-phase nature of our high-pass filter, it affects the phase rotation of every band. That means it must be convolved with each of the speaker bands. After that, we need to double-check that all the crossovers actually sum back into our ideal speaker magnitude response, and that the shape of the impulse response stays the same. This is indeed the case:
Note: if we look at the IRs of individual crossover bands, we see that their peaks do not align at the same position. This is totally correct because our model has a minimum-phase component:
Having this reference alignment will be very handy later for time-aligning the drivers. But before that, the drivers must be linearized in order to coerce them into the shape of the crossover bands we have defined. And before doing even that, let’s do a quick experiment.
What happens if we apply Acourate’s room/speaker correction to the IR of our ideal speaker? Interestingly, it “disagrees” that it is ideal and corrects it! How exactly? Acourate actually brings it to a minimum-phase behavior:
(The red trace is the phase of the ideal speaker after Acourate’s correction; the magnitude response stays the same.)
Why? The philosophy of Acourate is that speakers are minimum-phase devices. Plus, as we discussed before, minimum-phase filters exhibit no pre-ringing at all, by definition. One of the goals of the acouStep algorithm is to avoid pre-ringing.
The process of the speaker correction builds a filter which removes the excess phase. Typically, this excess phase comes from analog crossovers (in passive speakers) and speaker driver imperfections (we will see that later, once we start to combine linearized drivers). However, the linear-phase behavior at the end of the speaker frequency range in our ideal model is “excess phase” too! Acourate does not distinguish between “good” excess phase (like in our model’s case: it has excess phase because the model is linear-phase, but this pre-ringing is imperceptible) and “bad” excess phase (from minimum-phase crossovers, potentially audible); it just removes all of it for the direct speaker sound.
Should we worry about this? I don’t think so. First, we actually still use the ideal speaker model for defining crossovers. Second, when the speaker magnitude response is mostly flat, there is no big difference in the phase behavior between the linear-phase and minimum-phase model. Third, as follows from my brief analysis of the audio interface behavior, the phase behavior at the end of the frequency range may be severely affected by its brick-wall filters, thus we risk over-correcting.
Now, back to the main path. With the linear-phase crossover components developed, we can use them directly with the sinc-pulse linearization function of Acourate. Applying it to each driver of this system has its own intricacies, though.
The most straightforward is the woofer because it uses a classical sealed enclosure arrangement. In order to simplify microphone placement, I used a version of LXdesktop stripped of the full-range driver:
Although the full-range driver as a physical object certainly creates some reflections, due to its size they mostly affect mid- and high-frequency ranges which we taper with the crossover.
The full-range driver works without a baffle (as a dipole) and thus has backward radiation; measuring it in the near-field does not capture the dipole cancellation and thus does not reveal its actual far-field behavior. However, we don’t want to measure it from too far (the main listening position, MLP) because that will incorporate unrelated reflections. I experimented with the distance to ensure that the microphone is not very far from the driver, while the impulse response still looks similar to what I’m seeing from the MLP. Because of these constraints, the IR of the full-range driver looks more jagged; however, that is fine because Acourate applies some smoothing in the sinc-pulse linearization procedure anyway.
I have to say that the sinc-pulse driver linearization procedure is much more straightforward than my previous experience with Acourate V2 where I was applying the speaker correction to each driver and also manually removing phase deviations by means of all-pass filters. One important thing to remember is that the process of probing the speaker driver with a sinc pulse is much more sensitive to background noise. In order to get a measurement with a high SNR, the procedure needs to be repeated about 200 times (though this is not a major issue, since it runs fully automatically).
And for the subwoofer, I actually didn’t use the sinc-pulse linearization at all because the near-field sinc-pulse measurement has shown that it does not have many irregularities to compensate. So I decided simply to apply my raw crossover to it, and then iron out the magnitude response shape during the room correction procedure. After all, the low-frequency output in a room is dominated by room modes, which induce far greater magnitude swings than any native irregularities of the driver itself.
As the result of linearization, we have the impulse response of each driver brought as close as physically possible to its crossover band. What we can do at this stage is to see what happens if we align them in the same way as the IRs of our model’s crossovers and then sum them. This is what happens:
You can see that although it looks close to the shape of the ideal IR, there are some notable differences. First is that it is contaminated with reflections—this is unavoidable because the linearization was done in a home room, and Acourate does not try to fill up every possible dent in the magnitude response in order to avoid over-correcting.
The second difference is more interesting: we can see that the real speaker’s impulse has some pre-ringing ripple. This is not the result of bad time alignment—the impulses of real drivers are aligned as close as possible to their ideal speaker counterparts. The ripple is caused by non-ideal summation of pre-ringing from the drivers’ crossovers. Drivers simply can’t behave as ideal crossover band filters. Due to physical constraints, they have what is called “passband ripple”—unavoidable deviations from the target response.
Why this pre-ringing is not a problem: first, thanks to the initial choice of the crossovers, it falls under pre-masking thresholds; second, after we assemble the real speaker’s response, we can optimize it with acouStep. In fact, we can preview the result by applying Acourate’s macros to our intermediate IR as if we had measured it. Below is the optimized IR produced by running the “test convolution” with Acourate’s correction filter:
We can see that acouStep is capable of shifting the energy from the pre-ringing to post-ringing thus totally sweeping it under the rug of the post-masking of human hearing which has an even longer span than pre-masking. Finally, below are group delays of the ideal speaker, the uncorrected “real speaker”, and its corrected version:
As a reminder: the red trace is the “real” speaker, the green trace is the “ideal” speaker, and the brown is the corrected “real” speaker. Mostly we see that the variation of the group delay is reduced across the frequency range, and the group delay below 20 Hz is reduced (although this is not that important). Let’s also look at the phase:
The phase graph above shows once again that Acourate corrects the speaker to minimum phase—the phase is not flat in the high frequency region where our ideal speaker (it was used as the target here) has the roll-off.
With driver linearization complete, and having confirmed that our assembled speaker should mostly conform to the ideal speaker behavior, the next step is to “assemble” the output from the whole speaker by aligning individual driver outputs in time and adjusting their relative gains. I used to employ the sine-wave convolution approach for that; however, after considering the coherence (or, mostly, the lack of it) between the woofer and the subwoofer, I started to doubt that approach: aligning them at a single frequency point may be an over-optimization which makes summation in the adjacent regions not so ideal.
Learning from the experience of assembling the speaker response from the IRs of individual drivers, I realized that the same approach can be used for measurement of the drivers from the MLP. One difference is that we absolutely need to establish the timing reference. Typically, the driver at the highest frequency band is used for that because its impulse is the narrowest and “tallest” one. In order to split out the IRs of individual drivers within one single measurement, we can use the “impulse shift” method which is described both in M. Barnett’s book on Acourate and the free manual by Dr. Keith Wong available at the Acourate user forum.
The idea is that we delay the time anchor driver (in the case of the LXdesktop this is the full-range driver) by some known amount, for example 1000 samples or more. This makes the IR of the driver being aligned come before it. For example, for the woofer driver, the measured IR with the “impulse shift” looks like this:
The green trace is the IR of the crossover at its aligned position, shifted left by 1000 samples. This is our reference that we can align our real IR against. The strong peak to the right is the impulse of the full-range driver.
Because the woofer’s IR precedes the IR of the full-range driver, it is not obscured by the room reflections of the full-range driver’s pulse. Note that the IR of the woofer is relatively compact, thus the separation by 1000 samples is enough to see its main pulse part. The IR of the subwoofer is more smeared in time, and I had to use a 3000-sample offset instead.
You can see that the subwoofer’s impulse is so long that the delayed impulse of the full-range driver “rides” on it. Also note that there is a significant distance between where the subwoofer’s impulse needs to be and where it is before the alignment.
How do we perform the alignment against the reference? For the woofer’s IR this can be easily done “by eye” by looking at the IR peak (Acourate can also show its sample index, as the local maximum in the selected area). The subwoofer’s IR (especially the one from the actual physical subwoofer) may look more ambiguous. I ended up asking Claude to write me a MATLAB script for aligning IRs via cross-correlation. It also suggested other methods for subwoofer alignment such as GCC-PHAT, however their application actually resulted in poorer alignment than straightforward cross-correlation (which is not too surprising: the PHAT weighting whitens the spectrum, so for a band-limited signal like the subwoofer’s it gives equal weight to the out-of-band bins which contain only noise). But I assume this is where YMMV based on your actual room setup.
After time-aligning the drivers, I also performed level alignment by looking at magnitude responses with a frequency-dependent window (FDW) applied and adjusting the gain of filters.
The final check of the alignment is done by measuring the complete speaker. We can see that the IR of the actual speaker is indeed close to the ideal speaker, and exhibits the same pre-ringing issue that we already saw with our synthetic “real speaker”:
But we know that this can be corrected. In fact, because each driver would require a different passband ripple correction, it’s much easier to accept it first, and then remove it at the final stage.
As the last step, I performed the overall room/speaker correction in Acourate. I used the new “acouStep” option when making it, and the resulting step response looks rather nice, featuring a sharp onset and almost no pre-ringing:
(Note that here red and green are the left and the right speaker, respectively.) Did we achieve the desired linear-phase behavior? Yes—if we look at the phase and the group delay of the FDW-windowed response, we can see that they are flat, except for the low end, below 100 Hz where, as we know, group delay non-uniformity is imperceptible:
Recall that Acourate actually corrects toward minimum-phase behavior. This is why there is a phase roll-off at the high end. However, this is beyond my hearing range anyway. Also, as I mentioned in the beginning, if we try to correct that “imperfection,” we risk over-correcting.
As mentioned before, Acourate places the start of the impulse at sample position 6,000; thus, the baseline group delay is approximately 125 ms. The significant phase deviation near 100 Hz (and its corresponding group delay dip) is caused by a boundary reflection, as is the phase ripple near 300 Hz in the right channel. Trying to correct these boundary cancellations in DSP makes little sense.
Just to outline our approach once again, step by step:
And as a result we end up with a “linear phase” speaker (by industry standards) which is actually mixed-phase, but has linear-phase behavior where it matters—in the passband.
Having the sound system set up, I did some listening to albums that feature enveloping, tonally rich compositions with a non-trivial spatial layout. I was listening using the psychoacoustic correction for the center channel and the diffuse field, which essentially adds virtual center and surround speakers (although this effect is limited to the MLP only).
These are the tracks that I used:
I’m almost finished creating the new speaker setup; the only missing piece is the diffusers (arriving soon!), which should help remove some asymmetric reflections from the back wall. I think that this setup can match the symmetry and linearity of earspeakers, so that comparing them one-to-one becomes more straightforward.