I wonder what's really the limit in Q10a. For example companies will have huge recommendation models built, likely very automated without manually assigned labels. If they never ask for "“yes,” “no” for whether the natural Person is of Hispanic, Latino, or Spanish origin", but one of the automated recommendation dimensions is close to 1:1 match for that question, would they have to disclose that? (i.e. give statistics per-dimension)
This applies to the rest of Q10 too. 1000 top attribute values? Does "around vector (0,0,0.3,0.9,....)" count? Because barely anyone asks explicitly for the real user attributes anymore.
I think both assertions are wrong: Regulators still have no real conception of how any of this technology works, and you can watch the congressional hearings with Facebook for evidence that mostly they didn't really have a clue what to do once they had him in the room, or how exactly Facebook even makes money.
The court system only confirmed your rights to use webcrawlers within the last year, and webcrawlers have been around since the start of Google.
And as for the series of tubes guy, I'll give him the benefit of the doubt. He clearly still didn't have a good grasp on the subject, but I don't think he was being literal.
(for the record, I don’t mean to call out the fellow who actually made the series of tubes statement, it was just clear to me that even as an analogy it showed a lack of understanding of the basics of how the internet actually functioned, and that people in his position had no ongoing mandate to understand it.)
Here’s an example of what I mean by they have more understanding now:
I always thought it was a pretty good metaphor and didn't understand why he got so much grief for it. Breaking up messages into little chunks, sticking each chunk into one of those little capsules like they have at the bank, putting each capsule into a pneumatic tube, sending it down a series of tubes with switching stations along the way, and reconstructing all the chunks (which may have taken different paths through the series of tubes and may have arrived out of order) at the destination -- this is a good metaphor for the internet!
> If they never ask for "“yes,” “no” for whether the natural Person is of Hispanic, Latino, or Spanish origin", but one of the automated recommendation dimensions is close to 1:1 match for that question, would they have to disclose that? (i.e. give statistics per-dimension)
I'm going to assume yes, they would need to disclose that. That scenario sounds awfully similar to some of the terms the FTC uses when discussing the Fair Credit Reporting Act (FCRA):
> Under traditional credit scoring
models, companies compare known credit characteristics of a consumer—such as past late payments—with
historical data that shows how people with the same credit characteristics performed over time in meeting
their credit obligations. Similarly, predictive analytics products may compare a known characteristic of a
consumer to other consumers with the same characteristic to predict whether that consumer will meet his or
her credit obligations. The difference is that, rather than comparing a traditional credit characteristic, such
as debt payment history, these products may use non-traditional characteristics—such as a consumer’s zip
code, social media usage, or shopping history—to create a report about the creditworthiness of consumers
that share those non-traditional characteristics, which a company can then use to make decisions about
whether that consumer is a good credit risk. The standards applied to determine the applicability of the
FCRA in a Commission enforcement action, however, are the same.
> Only a fact-specific analysis will ultimately determine whether a practice is subject to or violates the
FCRA, and as such, companies should be mindful of the law when using big data analytics to make FCRA covered eligibility determinations.
Around here we've had a recent law mandating that any government-related decisions that use an algorithm must, upon request by a citizen, be explained in plain language, following all the steps in the algorithm.
I can't wait until they have to do that for a black box, neural network program !
It doesn't even require anything esoteric like a neutral network. This would already be very difficult for the federal reserve for instance, or the SEC. Or even the bureau of land management. Lots of government agencies crunch financial and other kinds of data to make decisions, which would be very difficult to describe in lay terms.
I wonder what's really the limit in Q10a. For example companies will have huge recommendation models built, likely very automated without manually assigned labels. If they never ask for "“yes,” “no” for whether the natural Person is of Hispanic, Latino, or Spanish origin", but one of the automated recommendation dimensions is close to 1:1 match for that question, would they have to disclose that? (i.e. give statistics per-dimension)
This applies to the rest of Q10 too. 1000 top attribute values? Does "around vector (0,0,0.3,0.9,....)" count? Because barely anyone asks explicitly for the real user attributes anymore.