Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Friday, 20 July 2012

'Informant incompetence'

I can't now remember where I heard the phrase 'informant incompetence', but it's a slightly cruel way of describing a perennial problem in linguistics (and presumably other disciplines too): when the people giving you linguistic data simply fail to understand what you want from them.

Tuesday, 29 May 2012

Double modal or double fluff?

Those of us who are interested in dialect syntax but don't make it their business to conduct experiments into it are always on the listen-out for interesting examples. You can't help it, after a while. On the Antiques Roadshow back in April, I heard one of the experts say this:
What date would that might have been?
He didn't stumble over it, it was very fluent production, so he either meant to say it or didn't notice what he'd said. But we seem to have a double modal construction here, something which is not found in Standard English and is attested but not common in certain dialects.

The modals are would and might, and if we put the sentence into a declarative form, you can see what the issue is:
That would might have been what date.
Either modal on its own is fine, but both together is not permitted in Standard English. As this is not part of my dialect I can't be sure that this particular combination is allowed in any dialect, but certainly two modal verbs can co-occur in many people's speech.

Not, however, in the antiques expert's speech, I'll bet. I would put money on this being a performance error, which went unnoticed because the fronting of the first modal would means that it's not adjacent to the second modal might. I would guess that he started out asking what date it would have been, and switched halfway through to asking what date it might have been, and the two met in the middle in a sticky mess. Perhaps the much higher frequency of would-questions than might-questions had some influence too (frequency estimation not based on any data or actual facts at all).

This kind of thing makes it so much harder to do dialect syntax through data collection. You might only have a few instances of double modal questions in hours of data, if you're working from interviews, and if a couple of them might be performance errors, how can you be sure of anything? This is why dialect syntacticians have to be cunning as a fox who's just been appointed Professor of Cunning at Oxford, and devise data collection methods that they think will cause people to use more double modals, but without telling them that they want them to use double modals. And getting people to say something in a certain way is really bloody hard. Normal people seem to have this quaint idea that what you say is more important than the way you say it.

Saturday, 11 February 2012

Correlations in linguistic data

Geoff Pullum at Language Log recently reluctantly (because it's not yet published) commented on a paper by a Yale economist, Keith Chen. In this paper, Chen argues that if your language has a grammatical future tense marker, you are less likely to save money, live healthily etc because the future seems like some other time, not to be worried about now. If your language uses present tense to refer to the future, you treat is an extension of the present and you'll be much more sensible about it. Pullum is guardedly sceptical about these claims, for reasons which you can read about yourself. 


He is also sceptical about this kind of claim (made based on correlations found in large amounts of data) because
I also worry that it is too easy to find correlations of this kind, and we don't have any idea just how easy until a concerted effort has been made to show that the spurious ones are not supportable. For example, if we took "has (vs. does not have) pharyngeal consonants", or "uses (vs. does not use) close front rounded vowels", would we find correlations there too? I have some colleagues here at the University of Edinburgh, within Simon Kirby's research group, who have run some informal experiments on the data Chen uses to see if dredging up spurious correlations of this kind is easy or hard, and so far they have found it jaw-droppingly easy.
He doesn't comment further on these experiments, but it reminded me of the talk Martin Haspelmath gave when our university's linguistics research centre opened a few years ago, and he told us about the World Atlas of Language Structures (WALS). After telling us what a wonderful, useful tool it is (and it is, I've found it invaluable), he ended on a note of caution. It's easy, he said, to find false correlations. For example, you can show a map of languages which have a different word for hand and arm or use the same word for both. That map shows that the languages that don't distinguish are, broadly speaking, around the warmer areas of the globe (yellow dots) and the ones that do distinguish are in colder areas (red dots):
(Map from WALS, feature 129A)
Now might one not hypothesise, asked Haspelmath, that this language fact is due to the climate? In colder countries the distinction is important, in that one wears items of clothing that cover only the hands (gloves), or sleeves that come down to the wrist. In warm countries, sleeves are not so long and gloves are not worn, so a separate word for hands never becomes necessary. A far-fetched example, but a lesson in not putting too much faith in correlations.