Posted by datakan 2 days ago
"Does the NSA collect any type of data at all on millions or hundreds of millions of Americans?"
"No, sir."
That case sure put Mark Russinovich on everyone’s radars, at least.
Good times.
As the saying goes, this is what radicalized me. If any one of us did anything close to this we would be treated like Aaron Swartz and given a 35+ year prison sentence.
But in the Sony case the DOJ didn't even threaten criminal charges.
> LG smart TVs do not continuously record or transmit users’ conversations.
"or" is ambiguous in the English language. A lawyer can rightly argue that they mean exclusive or and thus they are being technically truthful as long as they both continuously record and transmit users' conversations. They can resolve this ambiguity by stating each element as a independent sentence.
"continuously" means without end. A lawyer can rightly argue that as long as the recording can end in a single instance, then they are being technically truthful.
> Speech-to-text processing begins only if a user activates a voice interaction through a supported wake-word feature or by pressing the voice (or AI) button on the remote control.
"processing begins" makes no indication as to when it ends. A lawyer can rightly argue that as long as you use a wake-word a single time or press the voice button on the remote control a single time, they can begin processing and never stop. They can resolve this ambiguity by stating that they only process audio during a session and that session has a strict maximum duration.
Also, it makes no statement as to audio processing, only speech-to-text processing. A lawyer can rightly argue that as long as they do not convert the speech to text and just directly transmit the audio they are being truthful. They can resolve this ambiguity by removing the narrowly defined "speech-to-text" and changing it to "audio". Weird their lawyers made this so specific.
> Audio used for wake-word detection is processed locally on the TV and, if no wake word is detected, audio is not converted to text, stored, or transmitted.
Narrowly defined to be only "Audio used for wake-word detection". A lawyer can rightly argue that if they make a second copy of the audio that does not go to wake-word detection, then they can convert it to text, store, and transmit it. They can resolve this ambiguity by removing the narrowly defined "Audio used for wake-word detection" to just state that they do not convert to text, store, or transmit any audio outside of a session. Weird their lawyers made this so specific.
> Voice-recognition results and related technical logs may be generated as part of processing a voice command. These records are associated with specific voice interactions and do not indicate continuous recording of conversations occurring outside an active voice recognition session.
"These records ... do not indicate continuous recording" is not a denial that they are continuously recording. It merely states that it does not indicate continuous recording. A lawyer can rightly argue that as long as they have at least one record that is not associated with a continuous recording, then they are being truthful. No need to remove ambiguity here as it would be covered by the above fixes.
> Speech-recognition results may be used to support voice-related features but are not uploaded later when the TV is offline or when connectivity is restored.
"but are not uploaded later when the TV is offline or when connectivity is restored". Again, "or". Only mentions later, no statement about "now". Their lawyers can rightly argue that as long as they upload them immediately they are being truthful.
"uploaded later when the TV is offline" is illogical nonsense, how is it uploading when it is offline? Their lawyers can rightly argue that as long as they upload later when the TV is online and the TV never lost connectivity, then they are being truthful.
They can remove the ambiguity by stating that they never upload the speech-recognition results or only retain them until the voice-related feature has completed the task. Weird how their lawyers made this so specific.
> ACR uses audio fingerprinting technology using the TV’s internal audio processor (not a speaker) to identify content and does not collect screenshots, screen recordings, video recordings, voice recordings, or other audio recordings from the TV.
Again, "or". "uses" does not mean exclusively uses. "collect" only means ACR does not collect it. This does not indicate that screenshots, screen recordings, video recordings, voice recordings, or other audio recordings are not collected by other processes. It does not indicate that they do not use the resources collected by those other processes. They can remove the ambiguity by stating that ACR does not "use" these data sources and exclusively uses audio fingerprinting technology. Weird how their lawyers made this so specific.
Truly so odd how their lawyers make such precise, minute, and nuanced distinctions for their benefit, but leave everything else so ambiguous they can rightfully argue a tortured interpretation is technically truthful. If they were lawyers on behalf of the consumers, they would never accept such ambiguous language. Must be accidental.
Also they provided a statement earlier to Gizmodo: https://news.ycombinator.com/item?id=49628328
I am absolutely certain that ACR was enabled on our new TV by default, back some years ago. I am nearly certain their voice recognition, and advertising spying were enabled on our TV by default.
If none of this is truly "enabled by default" then this is a recent development.
I discovered this many years ago and did some deep dives, and unfortunately didn't publish anything about it. Lesson learned.
I really wish there was a better way to do a media center linux.
Kentucky passed HB 692 unanimously requiring smart TVs to ask permission if they are going to be spying on us. Ask your state legislature to follow their lead and pass some very popular consumer privacy protection.
So, was realme transcribing all conversations and uploading it? Can someone answer?
For example, let's say an ad network knows your and your friends' location, and knows you've met because you've been close together frequently, or are using the same wifi network. You talk about perfumes, then later one of them looks up perfumes on their phone. Google would usually have access to this much metadata, and will likely serve you perfume ads because of your friend who later looked it up.
This doesn't directly answer your question though. The answer would be - it's possible, but a lot is possible without that.
It's more likely that this is all due to external forces. Having your device on the same network, in close enough proximity to detect each other, for instance, combined with one of your friends having Googled the stuff you were talking about.
Or just basic data collection. There's the classic story about how supermarkets know you're pregnant before you do, just based on the groceries you buy (no, not pregnancy tests). Ad networks buy and collect a humongous amount of information about everyone and serve ultra specific targeted ads based on nothing more than pattern matching based on what the rest of humanity is doing.
It's impossible to rule out that the thing did run a 24/7 TTS service, but I think the depressing fact that ad networks don't need to listen to our conversations to know what we're talking about is more probable.
And your phone most likely already listens all the time to catch an "Hey, Google" (or something similar).
For example, last week I was complaining to my spouse about shoulder pain and how I was afraid I may need another surgery. YouTube has suddenly starting flooding my feed with videos about "shoulder exercises to avoid surgery." I hadn't searched any of those terms, purchased anything or visited an orthopedic, so I don't know what else could possibly explain it.
That’s the next possibility to eliminate.
I am positive I get ads based on the interests of other people the wifi where I live.
The definition of “process” is really important and is missing. LG could process all transcripts on their servers too and stay within a very wonky definition of “process” on the TV that excludes transcription and transmission.
So the TV gets advertised as "$$$, but basically $$ if you sign up to be surveilled and get a $/month check back." If the company sets $/month too low--or is too creepy about what they collect--people just won't opt-in.
But what happens if they promise big pay-backs but routinely cancel from their side? That's a problem, but at least it's a kind the market [0] can handle when it comes to reputation and bad reviews.
[0] "Markets are good, use them." -- "OK, fine, you must price this on a market." -- "NOT LIKE THAT!"