mariuscomper.ukRomână

Gereon, Cape Town, 2018

How late can a voice be and still be the same person?

You are close enough to see the mouth. The voice is still in it.

Before the mouth

The voice got there first

Hear it ahead of the mouth.

On the fitted curve, the 50% boundary is 131 milliseconds early.

After the mouth

The voice is late to the mouth

The voice is slipping later.

The other 50% boundary is 225 milliseconds late.

The width of a table

Two metres of air

The voice is coming back.

Two metres of air is 5.8 milliseconds. The page is adding that much delay to the sound now.

Step back and the voice can only get later. No distance makes it arrive early. An early voice has to be made in an edit or on a line.

The person in front of you

If someone is here

With their permission, film the person in front of you. Eight seconds, on this device, gone when you close the page.

Up to eight seconds. It stays on this device and disappears when you close the page.

The same air

Off the mouth again

The voice is leaving the mouth again.

Set again to 131 milliseconds early.

Across from you

Back across from you

The voice is coming back to the mouth.

At two metres the page adds 5.8 milliseconds. On the fitted curve, that point is still above the 50% line.

If you can sit where you see their mouth, the air has already added some delay to the sound. It can still be them. The study asked whether the voice and the mouth were in sync.

Method

The curve is from Conrey and Pisoni, 2006. Thirty-nine undergraduates began the task. Three had reversed the two buttons. Six answered synchronous on more than half the trials at every tested offset, out to a voice 300 milliseconds early and 500 milliseconds late, and no curve was fitted to them. The average is the other 30: 25 women and 5 men, aged 18 to 22.

They watched ten English words from one woman and pressed "in sync" or "not in sync". Headphones were at 70 dB. Delays ran from 300 milliseconds of voice lead to 500 milliseconds of voice lag, in steps of 33.33 milliseconds, because the pictures were 30 frames a second. A Gaussian curve was fitted to the averaged answers. The center is the mean point of synchrony. The edges are where that curve falls through half. The table gives that fit. The second number in each cell is the standard deviation printed beside it in the paper. The experiment measured perceived audiovisual synchrony, not whether listeners thought the voice and the face remained the same person. That is this page's framing.

Dixon and Spitz, 1980, are cited here from later accounts, not from their own pages, which were not available. Conrey and Pisoni list them as 18 listeners and give speech edges of 131 and 258 milliseconds, against 75 and 188 for a hammer hitting a nail. 258 milliseconds is 88.5 metres of the same air. Vroomen and Keetels describe the clips as starting together and then drifting apart at a constant 51 milliseconds a second, up to 500 milliseconds. Venezia, Thurman, Matchin, George and Hickok phrase the same procedure as steps of 51 milliseconds. Dixon and Spitz suggested the late allowance is learned from the slowness of sound. The tone and the circle, in the later study, carry a similar center, so that suggestion is not settled by these numbers.

Schwartz and Savariaux, 2014, recorded one French speaker repeating eight syllables six times. They argued that a mouth does not lead a voice by a fixed 150 milliseconds through ordinary speech. In the chained syllables, audible and visible events stayed between a voice 20 milliseconds early and a voice 70 milliseconds late. That describes those syllables.

The metres are a conversion, not a result from either study. Speed of sound at 20°C is taken as 343.2 metres a second, from 331.3 times the square root of 1 plus the temperature over 273.15. A swing of 10°C changes the two-metre delay by about 0.1 milliseconds. Humidity and altitude are left out. Two metres stands in for a nearby table. 47 milliseconds is 16.1 metres of that air. 225 milliseconds is 77.2 metres. The listeners were not standing there.

The faces are 24 frames a second, so the mouth itself changes about every 42 milliseconds. The delay of the table is smaller than one frame. You hear it as a shift in the sound. The control moves in tenths of a millisecond. That is finer than the study, and finer than the frames. A positive setting delays the sound. A negative setting does not play the sound early. It holds the picture back, using frames kept from the moment just before. Voice early means the sound is ahead of the mouth on the screen. The numbers are the settings the page asks for, not a delay measured at your speakers. Each clip loops, with a short fade at the join. The fade is ours.

Gereon was recorded by Commons contributor Bogreudell at Wikimania 2018 in Cape Town. The file is Wikitongues, CC BY-SA 4.0. This page uses 7.0 seconds, from 66.5 seconds in. Raluca was recorded by Nick Panzarella in Cluj-Napoca in 2017, also Wikitongues, CC BY-SA 4.0. This page uses 5.22 seconds, from 36.9 seconds in, and crops the black bars. Both files were transcribed before the trim. Eminescu died in 1889 and the verse is in the public domain. The trimmed, cropped, re-encoded clips are adaptations and are shared under CC BY-SA 4.0.

A recording you make stays on this device and is not sent anywhere. Echo cancellation, noise suppression and automatic gain are requested off. The page does not check what the browser applied. Sound stops while the camera is open, so the clip does not include the voice already playing. Closing the page drops the clip.

Fit to the averaged answers. The second number is the standard deviation from the paper.
CenterWidthVoice earlyVoice late
Speech47, 15357, 61131, 31225, 36
Tone and circle47, 43400, 66153, 46247, 61

The adapted clips are shared under CC BY-SA 4.0.

Early on the left. Inside the gold band, the fitted curve is above 50% synchronous responses.