Showing posts with label language. Show all posts
Showing posts with label language. Show all posts

Monday, July 20, 2015

Indo-European Languages and the Aryan Fallacy

In 1938, J.R.R. Tolkien (left) received a letter from a German publisher who proposed to publish a translation of The Hobbit. According to the laws under the Nazi government, they requested him to confirm that he was "of Aryan origin".

Tolkien was livid. He made his opinion clear in a related letter to his British publisher, where he said, "I have many Jewish friends, and should regret giving any colour to the notion that I subscribed to the wholly pernicious and unscientific race-doctrine."1

In his reply to the German publisher, he feigns innocence at first. "I regret that I am not clear as to what you intend by arisch. I am not of Aryan extraction: that is Indo-iranian; as far as I am aware none of my ancestors spoke Hindustani, Persian, Gypsy, or any related dialects."2

Tolkien had no truck at all with the Aryan Fallacy, but many of his contemporaries embraced it enthusiastically, including several other fantasy writers active at a similar time — very notably Robert E. Howard. So how did it come about?

The word Aryan correctly refers to a group of nomadic tribes that moved into northern India about three thousand years ago. It's also used for the linguistic family descended from Sanskrit, deriving ultimately from the language(s) spoken by these tribes. Lastly, and with extreme caution, it can be applied to the people who speak these modern languages (most of the languages of Pakistan, Bangladesh, Nepal, northern and central India, and the majority language of Sri Lanka). As long as it's understood that this in no way represents a racial identity.

The family, as Tolkien indicated, also includes Romani ("Gypsy"), its original speakers having migrated from north-west India. Ironically, in Hitler's time (i.e. before the mass migrations from the sub-continent after World War Two) the only substantial ethnic group in Europe who could legitimately call themselves Aryan were the Romani — who were persecuted by the Nazis.

The most obviously related group of languages is the Iranian family, together forming the Indo-Iranian group (Iran clearly derives from the same root as Aryan). Today, the Iranian group covers most languages of Iran and Afghanistan, together with Kurdish, but in earlier historical periods Iranian-speaking nomads lived on the steppes north of the Black and Caspian Seas: Scythians, Sarmatians, Alans etc.

In the late 18th century, linguists began to recognise a clear relationship between the earliest forms of Sanskrit, Greek and Latin, and postulated that they might all have derived from a common source. With this leverage to build on, the Germanic, Celtic, Slavonic and other language groups were added to the growing super-family, nowadays known as Indo-European. An extinct family of Indo-European languages, Tocharian, was even spoken in north-west China during the Han era.

In the 19th century, it was proposed that the whole family should be called Aryan, in the erroneous belief that this was the earliest name of a people speaking an Indo-European language and therefore the most likely to be the original. Though mistaken, this was originally innocuous, but a tide of racial supremacism gradually saw the label become more and more equated with the "German Race". Max Müller, one of its originators, was later scathing about colleagues who confused linguistic and racial characteristics, suggesting that "Aryan race, Aryan blood, Aryan eyes and hair" were as absurd as "a dolichocephalic dictionary or a brachycephalic grammar".3

The tide was against him, though, and the concept of the Aryan Master Race became more and more widespread, eventually finding its lunatic home in the warped mind of Adolf Hitler. In fact, there have been suggestions that the characteristic differences between the Germanic languages (including English) and all other Indo-European languages may have been the result of an unrelated people abandoning their own language and taking up a broken form of Indo-European. In which case, the Germans have even less right to the name Aryan than most other Indo-European speakers.

The clear pattern of divergence among the languages through various periods has always suggested that it should be possible to trace them all back to a time when a single language, from which they are all descended, was spoken by a community in a relatively small area. The favourite explanation nowadays is the steppes of south-eastern Europe somewhere between 4500 and 2500 BC, but other propositions have ranged from the eccentric (such as the North Pole) to the more plausible (such as Anatolia or the north-western European plains). The image to the left shows one proposed model of expansion from the homeland.

It shouldn't be assumed, though, that these original Indo-Europeans represented a race, or even a "people" as we'd understand it. Archaeological evidence suggests that the communities who probably spoke Indo-European were actually made up of several distinct groups — a people who'd come down from the north, another who'd come up from the Mediterranean, and possibly others from Central Asia or the Caucasus — who can all be seen from their remains to have been very different physical types.

Nor are they likely to have had much in common, including a political or tribal structure, except the language which allowed ideas and customs to spread. They constituted a culture area, but no more.

In terms of the later spread of Indo-European, too, we can't assume any racial connection. It's not unusual for peoples to adopt the language of either the dominant or the "cool" culture, and this often means communities speaking a language aren't racially connected with the those who spoke the language's distant ancestor. If we didn't accept this, we'd have a hard time trying to explain how an English-speaking African American could have derived from the language's source in north-western Europe.

The subsequent evolution of the languages is likewise anything but straightforward. Linguists use the family tree model, (eg Latin is the "parent" of French, Spanish, Italian etc) and this is useful, and a gorgeously artistic interpretation of which by Minna Sundberg is shown right. But the influence of other, often unrelated languages is important too. This can be down to extensive borrowing of vocabulary for social or political reasons, such as how English, fundamentally a Western Germanic language, has a vocabulary heavily derived from Latin.4 It can also be explained, though, by the wave theory of language change.

The wave theory is a model which suggests that specific changes, whether a sound-change such as a final t changing to s or a grammatical change such as the development of grammatical gender, spreads from an epicentre and affects both related and unrelated languages. The next major change will have a different epicentre, and/or the speakers will have moved, so it won't be the same set of languages affected.

Ultimately, this will create a patchwork, in which it can be difficult to reconcile a language's position on its family tree with apparent similarities to more distantly related (or unrelated) ones. It's inconvenient, but we're talking about human behaviour. What do you expect?

So where did the original Indo-European language come from? If it was being spoken a mere five or six thousand years ago, it obviously can't have sprung from nowhere. The reason it's recognised as the "original" is that it marks the latest point that all Indo-European languages can be traced back to, but its history must have gone back a long way up an unknown family tree.

Various propositions have been made as to what Indo-European might ultimately be related to, the most widespread being a super-family known as Nostratic. At its most ambitious, this hypothesis includes Uralic, Altaic, Kartvelian, Afroasiatic, Elamo-Dravidian and Eskimo-Aleut, along with a few others, though some more cautious proponents restrict it to the first three plus Indo-European.5

There's very little evidence for hypotheses such as this, and what has been put forward is strongly disputed. The problem is that it becomes progressively harder to be sure of connections between languages the further back the proposed connection is. It's probable that some, at least, of these connections are correct, but nothing's likely to ever go beyond speculation.

To speculate, though, how far might this process actually go? Recent evidence suggests an origin for language considerably further back than was believed even a decade or two ago, certainly when our ancestors were living in a fairly small area of Africa.6 Maybe the invention of human language really was a single event, and we're all speaking variants of the same language.

The Aryan Fallacy should have died in the bunker with the Führer, but unfortunately bigots aren't famous for their intellectual rigour, and there are still white supremacist morons who use it as a keystone for their disgusting creeds.

So take a leaf out of Tolkien's book. Next time a white supremacist proudly claims to be Aryan, point out that they're actually claiming to be Indian. Or even Iranian. You won't change their views, but see how they like it.

 

1 The Letters of J.R.R. Tolkien, ed. Humphrey Carpenter, 1981, p.37

2 Ibid. p.37

3 Quoted in In Search of the Indo-Europeans, J.P. Mallory, 1989, p.269

4 That the roots of English are Germanic can be illustrated by taking a passage of average English and identifying where each word came from. It's likely that more would be Latin-derived than Germanic-derived; but if you were to redo the exercise counting each word each time it's used, you'd find the count overwhelmingly Germanic. This is because the most common, infrastructure words (the, and, to, for and the like) are direct descendants of the words that came over with the Anglo-Saxons.

5 Uralic includes languages like Finnish and Hungarian; Altaic includes the Turkic languages, Mongolian, probably Korean and possibly Japanese; Kartvelian is a small group whose main member is Georgian; Afroasiatic is a huge group, ranging from Hebrew and Arabic to Hausa, the main language of northern Nigeria, and taking in ancient Egyptian along the way; Dravidian was the main language group in India before the coming of the Aryans, and is still dominant in the south; Eskimo-Aleut — yes, I know Eskimo is non gratis, but there's actually no other word for the whole linguistic group, the Inuit actually being just the largest ethnic group.

6 Remains have indicated that Neanderthals shared the same deformity of the larynx that allows us to manipulate complex sounds, suggesting that the mutation took place at least 150,000 years ago. On the "use it or lose it" principle of evolution, it's difficult to interpret this any way other than our ancestors already using language from that point.
 
Images reproduced under Creative Commons licence:
 
Tolkien: Proyectolkien
Dancers of the Shuvani Romani Kumpania: James Niland
Indo-European expansion according to the Kurgan hypothesis: Dbachmann
Minna Sundberg Language Tree: Tom Wigley
 
 
 

Wednesday, May 6, 2015

Pronunciation of Names

In the note at the end of his 1935 epic fantasy novel Mistress of Mistresses, E. R. Eddison suggests: "Proper names the reader will no doubt pronounce as he chooses. But perhaps, to please me…" Give or take the odd gender-specific pronoun, I'd say the same about the names I use in my stories. It's really not a big deal if you don't pronounce them as I do, but if you're anything like me, you might be interested in how they're "supposed" to sound.

No, I'm not going to give a complete list of every name I've included in a story — that would be a very long article indeed — but I'll try to give an idea of what I understand by the letters I type. The first thing to bear in mind is — forget all about English. Well, unless I'm actually using English-style names, of course. English pronunciation is among the most idiosyncratic in the world, and in general it's better to think in terms of Greek or Latin for my letter values.

The second thing is that there are no silent letters. None. Not a single one, unless you count the h in combinations like th and sh. If a name ends in e, that e should be sounded, rather than modifying a previous vowel. Double letters are pronounced double. And, if a name starts with two consonants that don't go together in English, no wussing out (as we usually do with Greek-derived words, for instance). It's not rocket science. Little children all over the world have no problem learning to pronounce words that start with ks, tl, nd or mb.

The stress on a name is harder to generalise about. My names come from many different (theoretical) languages with different (theoretical) rules on stress. If in doubt, though, I tend to default to having the stress on the second-last if the last is a strong syllable (eg Renon = RE-non) or the third-last if the following syllables are weak (eg Caurien = COW-ree-en).

·         a — either as in cat or calm, never as in cake.
·         ai — as in aisle.
·         au — as pronounced in German, like the vowel in cow.
·         b — as in big.
·         bh — halfway between b and v. If that's too hard, pronounce it as v.
·         c — always as in cat, never as in ceiling (whatever follows it). *
·         ch — usually as in loch (in a proper Scottish accent). However, I'm a bit inconsistent with this, and occasionally at the start of a name it's as in church (eg Chenda from At An Uncertain Hour).
·         d — as in dog.
·         dh — like the th in this, as opposed to in think.
·         e — as in get, never as in cede.
·         ei — like the vowel in day.
·         eu — difficult to describe to everyone, as it's a sound Americans generally refuse to pronounce. As in French tu, or the way I'd pronounce few (fyoo) — not as in moon.
·         f — as in fish, never as in of.
·         g — as in get, never as in gin.
·         gh — sounds a bit like gargling.
·         h — always pronounced, wherever it occurs, unless it's part of th, sh etc.
·         i — either as in pin or like the vowel in been, never as in fine.
·         j — as in jam.
·         jh — like the middle consonant of treasure. This is pronounced exactly the same as zh — just an indulgence for variation.
·         k — as in kit. *
·         kh — same as the hard ch.
·         l — as in lid.
·         lh — like the Welsh ll. You can all pronounce that, can't you?
·         m — as in meet.
·         n — as in new.
·         ng — always without sounding the g. For my pronunciation, as in sing rather than finger, but not everyone differentiates these.
·         o — either as in cop or as in cope.
·         oi — as in boil.
·         ou — like the vowel in moon.
·         p — as in put.
·         ph — as in philosopher (both times).
·         q — similar to gh.
·         qu — as in quiet.
·         r — this is a difficult one. Different languages/accents pronounce r as anything from grating at the back of the throat to trilling, and it would depend on the character which is correct. I pronounce it from the front of the mouth but without a trill, but whatever's natural to you is OK. When it follows a vowel, it's sounded as well as modifying the vowel (ar as in car, er as in Ernest, ir as in fear, or as in form, ur as in cure).
·         rh — a more emphatic version of r (sorry, difficult to describe it).
·         s — as in sing, never as in rose.
·         sh — as in shout.
·         t — as in take (there are several different ways of pronouncing t as well — and d for that matter — use the one most natural to you).
·         th — as in think (see dh).
·         u — as in dispute, never as in dumb.
·         v — as in vague.
·         w — as in warrior.
·         x — as in axe. Always. Even at the start of a word.
·         y — as a vowel, like the indeterminate vowel represented by ə. If it starts a word, followed by a vowel, as in yellow.
·         z — as in zoo.
·         zh — the same as jh.
·         apostrophe — no, whatever you've been told, the apostrophe isn't just a decoration. When a non-grammatical apostrophe appears in my names, it represents the glottal stop, which can perhaps be best described as a cross between a gulp and a hesitation. It's not a standard sound in English (although it appears in some accents) but in many languages it's as much a letter as A or B. A good place to hear the glottal stop used is the TV show Stargate SG1, where several alien names, including the character Teal'c, have one in them when pronounced correctly.

* The keen-eyed among you may have noticed that c and k are identical. In fact, I arbitrarily use them to represent respectively the unaspirated and aspirated versions of the sounds. In many languages, aspiration is a vital distinction between sounds, but in English the two forms are used interchangeably. If you try pronouncing the words car and scar naturally (not easy when you're thinking about it) you'll probably find that the c in car is accompanied by a puff of air, whereas the c in scar isn't. That puff is aspiration. So c should be pronounced without aspiration and k with.

Or, if that's too complicated, just pronounce them both the same. Like Eddison said, as you choose.

Thursday, April 9, 2015

The History of the Alphabet

When people are asked about the greatest design achievements of all time, they might mention anything from the Periodic Table to the Tube Map. Those are certainly great examples, but no-one ever seems to bring up perhaps the greatest design in human history: the alphabet. Yet the alphabet is behind everything our civilisation has achieved in writing, from the works of Shakespeare to the last text you sent. Any of us who are writers rely entirely on the alphabet.

There was writing long before the alphabet, and there are still important writing systems around the world that have no connection to it. Widespread systems range from the ideographic Chinese characters to the semi-alphabetic Devanagari in India and beyond 1, as well as more localised systems, such as Sequoia's wonderful Cherokee script. Nevertheless, Latin script is by many orders of magnitude the most common system in the world, while two of the next four most common (Arabic and Cyrillic) are also descended from the original alphabet.

The earliest writing systems were probably ideographic, as Chinese characters still are. This means that the symbols indicate a concept, rather than a spoken word, a technique we use occasionally in the West. For example, "2" means exactly the same in every language, regardless of whether it's pronounced two, dos, zwei etc.

Chinese proves that ideographic writing can produce everything from great literature to great record-keeping, but its drawbacks can be illustrated by the history of printing. The Chinese didn't take to movable type, as the Europeans did, even though they had printing long before Europe. It wasn't that they hadn't come up with the idea of movable type, but it was impractical for the simple reason that printers would have needed to have been surrounded by thousands of different characters.

Many early writing systems, on the other hand, used syllabic scripts. This means, for instance, ba, be, bi, bo, bu and by would each be represented by a separate symbol, and words word be written simply with two or three of these. It was more straightforward than ideograms, but still required a hundred or more symbols to be learnt before you could read or write it.

No-one knows exactly who came up with the alphabet 2, nor exactly when, but it seems to have been invented by the Semitic 3 peoples of late Bronze-Age Canaan — roughly what's now Israel, Palestine and Lebanon. One hypothesis suggests that it was begun by a Semitic tribe in Egypt, simplifying the hieroglyphic system. Inevitably, this has prompted speculation that Moses was responsible, but that's a pretty long shot. Even if the breakthrough did happen in Egypt, there were many Semitic tribes living there during the relevant period.

Whatever its exact origins, the invention transformed writing in Canaan and beyond. The decision to reduce the symbols so that only those for each consonant were used had the advantage that there were now only a couple of dozen to learn, making reading and writing easier for non-specialists to master.

It came at a price, though. The Canaanite alphabet, whose closest modern descendent is the Hebrew system, didn't have any way of indicating which vowels to include in the words. If English were written this way, it would be impossible to tell whether bd meant bad, bed or bud — or, for that matter, bide or abode. All reading would be like the final round of Only Connect. 4

Still, the plusses must have outweighed the minuses, because the alphabet not only thrived but spread, most importantly in two directions. For one thing, it eventually formed the basis of the Arabic script, which is now one of the most widespread writing systems in the world. Arabic script developed out of the system used by the Nabataeans, a northern Arabian culture that flourished in what's now Jordan and the surrounding areas — their most important centre was Petra, the "rose-red city, half as old as time". The Nabataeans borrowed the alphabet from Syria, where it had spread from Canaan, and passed it on to the Arabian Peninsula.

More relevantly to the English-speaking world, the alphabet also spread west. The seafaring Canaanites from ports like Tyre and Sidon were known as Phoenicians — from the Greek word for purple, because their speciality was the insanely expensive purple dye everyone wanted — and they settled and traded all over the Mediterranean and beyond.

Somewhere around 800 BC, the Greeks encountered the Phoenician alphabet. The original Greek syllabic script (Linear B) had been lost in the dark age that followed the fall of the Palaces, and the Greeks took up this new idea with enthusiasm. However, aware of its shortcomings, they came up with the crucial idea of turning some of the letters they didn't need into vowels.

The original Greek alphabet wasn't quite the one used today. Over the next couple of centuries, they dropped a few letters and added others. One of the best-known Greek letters, omega, was a late addition, which is why it comes last. Before that process started, though, the great Italian civilisation of the Etruscans adopted the primitive Greek alphabet, and through them it came to a small city-state called Rome.

The Roman alphabet preserved letters lost in later Greek, such as F and Q, but it had its own problems. The Etruscans hadn't needed a G (originally the third letter, as in the Greek gamma) and took to pronouncing it the same as K, creating the modern C/K duplication. The Romans, however, did need a G and so converted the seventh letter, properly the Z, into their G.

The Roman alphabet expanded as they added letters needed to write foreign words. The Z was reinstated, but put at the end, and the Emperor Claudius (of "I" fame) invented the letter Y — which, by the way, is really a vowel occasionally pronounced as a consonant, not the other way round, whatever your teachers might have told you.

As the Roman Empire spread, so did both the Greek and Roman versions of the alphabet, Greek spawning various other forms, such as the Armenian and possibly Georgian scripts. When it was necessary to translate the Bible for the newly converted Slavs, Saints Cyril and Methodius (allegedly) came up with a new alphabet, which combined Greek letters with other symbols for Slavic sounds the Greek alphabet didn't cover. There's been considerable scholarly debate over whether these were derived from one or more other writing system, or whether they were invented. Whichever is true, varieties of the Cyrillic alphabet are now used throughout much of eastern Europe and a good deal of Asia, including most of the languages from the former Soviet Union.

Many of the letters have been written in several different ways even by the same people, and different forms were passed on as the alphabet spread. Both the Greek and Cyrillic versions include a number of "false friends", such as perhaps the most famous Cyrillic acronym, CCCP. This is the Russian name for the old USSR, but in fact C is the Cyrillic letter for S (derived from a form sometimes used in early Greek) and P is the R in both alphabets. CCCP should actually be pronounced SSSR.

Meanwhile, the Roman alphabet spread throughout western and central Europe, though like Cyrillic it was adapted to the needs of different languages. Old and Middle English, for instance, had four letters that we don't use now, including the þ, representing th. This was later often written lazily as y — hence all those "Ye Olde Tea Shoppes".

On the other hand, there were modern letters missing. Until the 17th century, I/J and U/V were each considered no more than different ways of writing the same letter. 5 The Romans had pronounced the consonant form of U/V like our W, but this gradually changed, both in Church Latin and the vernacular languages, to V as in Victor, so the W was invented to replace it.

By the 18th century, the 26-letter alphabet we know now was in place, although some languages (in Scandinavia, for instance) still use extra letters, and different descendants of the original Semitic alphabet are used all over the world. It may change again, of course, if it needs to. Just like all those other design classics — including the Periodic Table and the Tube Map — it has adaptability built in to accommodate change. One thing is certain, though — there's some long-dead Canaanite who deserves to be picking up a hell of a lot of awards.

 
1 One theory suggests that Devanagari is ultimately descended from the Semitic alphabet, while another insists that it's indigenous to India. The jury's out.

2 Technically, scholars of writing systems classify the Semitic system as an abjad, rather than an alphabet, since it doesn't use vowels. However, as we'll see, there's a direct lineal descent to the alphabet we use today.

3 Semitic is usually used today as synonymous with Jewish, but it actually refers to a group of languages (and, to a lesser extent, the peoples who've spoken them) which includes Hebrew, Aramaic, Assyrian, Babylonian, Arabic, the main languages of Ethiopia, and even Maltese.

4 For anyone not familiar with this fine quiz show, the final round involves two teams racing to be the first to recognise phrases from the consonants only.

5 Which makes a nonsense of the "Name of God" scene in Indiana Jones and the Last Crusade. I and J were the same letter in Latin, as they would have been when the trap in the film was first set.
 
Image courtesy of Tom Magllery, Creative Commons licence
 

Sunday, November 9, 2014

Heap Many Moons - The Way Primitive Peoples Don't Really Talk

We've all encountered old-fashioned idea of how "primitive" peoples speak: "Heap many moons we hunt buffalo," or some such nonsense.  After all, primitive people must speak a primitive language, right?

Of course not.  Even setting aside the question of how primitive the cultures in question are, there's really no such thing as a primitive language.*  Every language expresses precisely what its speakers need to express.  A language spoken by a nomadic tribe of herders and hunters might not have technical words for scientific or sociological concepts, but it'll have vocabulary that enables people to talk about a whole range of concepts, emotions and relationships that are important to them — many which can't be clearly expressed in English or other majority western languages.

The "heap many moons" style of speech is a very simple example of pidgin, a type of language that arises when people from different cultures need to communicate on a basic level.  It's usually for trade, but pidgins can be heard whenever Britons or Americans abroad are trying to make themselves understood to "foreigners who have the nerve not to speak English".

Where the contact is regular and long term, a pidgin can develop regular rules and vocabulary, and eventually, given the right circumstances, children might start growing up speaking nothing else.  At that point, it goes through a metamorphosis into a creole, a language that's flexible and rich enough for its speakers to say whatever they need to.  There are creolised languages from the Caribbean to the Pacific, and many are elegant and expressive.

All languages intended as the primary means of expression for a people are tailored exactly to what that people needs.  Whether or not the Inuit really have fifty words for snow,** they can certainly talk about snow in a lot more detail than a people whose language has evolved on the equator.

What a language does or doesn't have inevitably reflects what matters to the society.  Many of the Australian Aboriginal languages didn't have counting systems at the time Europeans first arrived.  This wasn't stupidity — if you rarely see more than a handful of any given object, including people, why would you need to count?  As soon as the concept was introduced to them, it took a remarkably short time for this gap to be filled.

On the other hand, many of them have degrees of sophistication in their grammar that European languages can't match.  English, for instance, has one way of expressing the first person plural pronoun — we (with the variants us, our and ours).  Some languages, though, (including Old English) distinguish between whether you're saying I and you or I and they.

This might seem strange to those of us whose languages have done without it, but it's actually a distinction between two very different concepts.  If you tell someone "We're meeting at eight o'clock," you might mean "We're meeting — you can make it, right?" or "Us lot — we're meeting up.  Just saying."  In English, we have to rely on tone and context to make it clear which we're saying, but it can be a useful distinction to make.

On the other hand, in many of those Aboriginal languages that didn't have counting systems there might be up to a dozen different ways of saying we, depending on exactly who the other person is, whether they're related to the speaker, whether or not they have the same Dreaming.  In these societies, it's vitally important to clarify these issues, and the languages have developed incredibly complex grammar to accommodate that need.

A similar, though simpler, concept that occurs in many European languages is the distinction between the familiar and formal versions of the second person pronoun.  This will be known to anyone who's learnt French, German or Spanish.  In French, for instance, you'd address a close friend or family member as tu and a more casual acquaintance or stranger as vous, while  German has equivalents for both the singular and plural forms.

English used to make this distinction, too, using thou and you, but thou has died out, except in a few dialects.  There are various theories for why this should have happened, but the effect (if not the reason) is that English-speakers don't have any of the complex social niceties needed to use this particular grammatical form.

The history of language is littered with abandoned grammar that once expressed vital concepts.  Early forms of the Indo-European language family, which includes almost all European languages,*** had not only a singular and plural, but also a dual number.  This may originally have been used to express any two things, though by the time it reached classical Greek it was only used for specific pairs: the eyes, the ears, egg and bacon, Simon and Garfunkel and so on.

On the other hand, the abandonment of the dual may reflect a fundamental change in how we view the relationships between things.  For us, there's an obvious difference between one and all other numbers, and that forms an essential part of the patterning of our minds.  There's a mathematical justification, of course, since one really does behave in ways that are different from every other number.  On the other hand, two is also a unique number, the only even prime, and viewing doubleness as a thing apart in the same way as singleness may have been integral to how those societies saw the world.  Perhaps it explains why triple deities are so common, if three was the first plural number.

What a language does or doesn't include can have an enormous effect on a society.  To return to classical Greek, there was a simple but far-reaching linguistic habit among the Greeks.  The language had two little words (men and de) that could each be slipped in as the second word of a clause to set it up in opposition to another clause or sentence.  You could roughly translate men as "on the one hand" and de as "on the other hand", but such little words could be used without the clunkiness of the English phrases.

It's unlikely to be a coincidence that the language which adopted this structure was spoken by a people who essentially introduced philosophy and logic to the west.****  Whether the structure nurtured a logical frame of mind or the impulse for logic created the structure (or a bit of both), a naturally dialectical language was perfectly adapted for Socrates, Plato and the rest to debate philosophy.

From subtle relationships to logic to high-tech (or even talking about snow), all languages are rich and expressive in the concepts their speakers care about — and, as long as they're human, that will certainly include a full suite of emotions and imagination.

If the hero(ine) in your story meets a primitive tribe, by all means show communication difficulties between them, but don't make the mistake of believing that really is how they speak or think, any more than your hero does.  They're probably too busy gossiping about the stranger who doesn't know the first thing about their way of life to bother with all those heap many moons.

 
* Not among any known human society, at least.  Various species of animals may have very simple languages — prairie dogs, for instance, appear to have a vocabulary of a few dozen words to tell each other about food and danger — and these would qualify as primitive languages.

** I'm fairly sure I've come across a debunking of that, but I'm not certain.

*** Except for Finnish, Estonian, Lappish, Hungarian, Turkish, Maltese, Basque and some minority languages in Russia.

**** It's often said that the Greeks "invented" philosophy.  Of course, the Chinese and Indians also "invented" it, and no doubt other cultures did too, but however dubious that claim might be, the Greeks certainly originated the western tradition of philosophy.

Sunday, March 30, 2014

How Did This Word Come to Mean That?

English is a funny old language.  In terms of formal linguistic classification, it belongs to the Low German branch of the Western section of the Germanic sub-division of the Indo-European family, but that's largely about where the language originally came from.  You can tell that from the words that don't tend to change much — the, and, but, to, for, one, two, three and the rest — but the vocabulary we use comes from all over. 

The original Anglo-Saxon speakers preserved many words from the earlier Celtic languages, just as they almost certainly preserved words from still earlier languages whose names we don't know.  Invasion, occupation and settlements gave us a huge shot of vocabulary from French and the Scandinavian languages, while the later rise of learning spawned numerous words derived from Latin and Greek.

In the past few centuries, English-speakers have conquered and colonised all over the world.  Besides exporting English, we've also imported vocabulary from many of these places — India (eg bungalow, pyjamas), Australia (eg kangaroo, boomerang) and the Americas (eg potato, wigwam).  And some words have just crept in randomly over the centuries, such as algebra from Arabic and robot from Czech.

Even in a ragbag language like English, some words have truly bizarre origins, and I thought I'd give a few examples.

1. Down


It's such a simple word, but extremely versatile, used as adverb, preposition, adjective, noun and verb.  It was originally a noun, though, and one of those words the Anglo-Saxons stole from the Celts.  The word dun actually meant a hill, and it still survives in the plural as downs, especially referring to the chalk hills in southern England.

From this derived adune, literally meaning from the hill — in other words, towards a lower level.  This is sometimes found in older or archaic language as adown, but it was soon shortened to its current form.

This was originally a preposition (He walked down the stairs) or an adverb (She put it down), but in modern English it can also be used as an adjective (He took the down escalator), a verb (The workers downed tools) or, coming full circle, a noun (She weathered the ups and downs of life).  All from an ancient word for a hill.

2. Item


The word item is actually Latin for also.  Its modern English use came from an old system, often found in Shakespeare's plays, for instance, of making lists.  A list would begin imprimis (firstly), and then each subsequent thing on the list would be preceded by item (also).

Over time, it became so common to write or reel off lists in this way that the word came to be seen as merely signifying the different "things" on the list.  Since there was no word for this at the time, they came to be known as items, and the meaning has since extended so it can refer to any discrete object or concept that might be (but isn't necessarily) part of a list.

3. Check


This is perhaps the strangest of all.  The word check or words closely derived from it can mean to stop something, to make sure things are OK, a pattern of squares or a promissory note from a bank, and you'd be forgiven for assuming some of these, at least, are unrelated homophones.

In fact, every single meaning of the word derives ultimately from the Persian word for king, shah.  This is normally pronounced in English without a final consonant, but I'd guess (I'm not a Persian speaker) the h should actually be sounded, giving something that could be distorted into check.

It was introduced to western Europe through chess.  Several chess terms derive from Persian (the rook, for instance, is a chariot) and you call out check to indicate you're attacking your opponent's king (effectively look to your king).  Checkmate means the king is dead.

As any chess-player knows, if you're put in check you have to suspend all your cunning plans to get out of check, so the word came to be used to mean stopping someone from completing what they're doing, such as checking an attack.  From that, it turned into checking yourself — looking before you leap — and then to mean investigating that something was as it should be.  Finally, in the US it's come to signify the mark known as a tick in the UK to show that something's been checked.

In the meantime, the word became attacked to the chess board, which became known as a checker or chequer board (the game of draughts, played on the same board, is known as checkers in the US).  From this, any pattern of alternating-coloured squares came to be referred to as a check or checkered pattern, and this is used metaphorically now, such as talking about someone's checkered past.

The chequer board, though, wasn't only used to play games on.  It was also the main form of abacus in mediaeval Europe, and its ubiquitous use led to any counting-house being referred to as an exchequer.  This included the government's financial department, which is why the chief financial minister in the UK is called the Chancellor of the Exchequer.  Banks, which were common in the Islamic world and brought to Europe by returning crusaders, had their exchequers too, and the promissory notes they issued were referred to as cheques.  In the US, they're checks, and this can also refer to a bill for payment, especially in a restaurant.

So, whether you're making an inventory, wearing a gingham dress or writing a note to transfer money from your bank account — not to mention getting the upper hand at chess — you're actually invoking the Kings of Persia.  Even if you don't know it.

Saturday, September 14, 2013

Sisters & Cousins & Aunts: Language Families for Fantasy Writers

Last year, I posted Fantasy Languages for Dummies here, where I outlined some of the basic issues to think about when inventing words and names for an imaginary language.  Judging from the number of hits, it seemed to be something that interested a lot of people, and I thought I'd try something a bit more advanced, for those who want a little more out of their fantasy languages.

Most fantasy writers who create a secondary world invent imaginary languages to some extent, even if it's only a handful of names.  Some, of course, stick strictly to real-world languages, but most have something invented.  In many cases, these don't offer any consistent sense of phonology or morphology, but there's usually a hint of it, even if it's only the Burroughs Universal Constant (that, wherever you go in the universe, female names always end in a).

Writers who take a genuine interest in language and naming, though, might put a lot more thought into the matter, creating names that follow similar linguistic forms when they come from the same culture and distinct forms when they don't.  Well and good; but few fantasy writers (unless they happen to be linguistically orientated professors of English from Oxford) seem to consider how their various languages relate to one another.

Languages don't exist in isolation.  Well, OK, some do, like Basque or Burushaski, but they're exceptions.  We'll come to that later.  Most languages, though, are grouped into a hierarchy of families, super-families, super-super-families etc. in structures so similar to biological taxonomy that the same terms are often used, though not consistently.  English, for instance, is a Low German language, belonging to the West Germanic division of the Germanic family, part of the great Indo-European phylum.

What exactly does that mean, though?  Everyone knows that English is essentially a rag-bag language that contains French, Latin, Celtic and Greek, as well as words from almost every part of the world the British Empire ever came into contact with.

There's an easy experiment that can show what the classifications mean.  Well, easy to imagine and explain: not quite so easy to do.  Take an average passage of English (not too scholarly or technical, not too monosyllabic) and list every word in it.  Then look up the origins of those words (a good dictionary would give that) and count how many words derive from each source.

What you're likely to find is that the great majority are either Germanic, whether from Anglo-Saxon or Scandinavian, or Latin, whether directly or via French.  There'll be a smattering of Celtic and Greek, together with odd words from further afield.  The balance of Latin and Germanic will probably be roughly even, maybe with more Latin words. 

So why isn't English counted as a Latin language?

Now repeat the experiment, but counting each word each time it occurs, and you'll find a dramatic change, with the vast majority derived directly from Anglo-Saxon.  This is because the Anglo-Saxon vocabulary of English includes the most common words, repeated over and over: the, and, to, for and so on.

This is what linguists mean when they classify a language as belonging to a family, and it's easy enough to see why.  Vocabulary changes all the time in a language.  It's not hard to imagine foreign words for, say, house or table becoming fashionable and eventually replacing the original words, but why would anyone create a different word for the?  These words do change, of course, but very slowly and usually only in ways that can be easily followed.  That makes them a breadcrumb trail back to where the language has come from.

Languages drift apart when two groups of speakers don't often interact.  They borrow different foreign words and coin different words for new concepts, but their pronunciation often diverges quite radically.  Pronunciation is changing all the time.  It's true what the old people say: youngsters don't speak the way we did at their age.  Of course not; and, over the centuries, it can become completely different, till a linguist comes along and explains how it's really all related.

There's a saying I remember hearing long ago: when comparing languages, vowels count for nothing, and consonants for very little.  That's not quite true, but you have to understand how the different letters are formed to understand why one can change into another.  One of the most famous sets of changes is known as Grimm's Law (yes those Grimms — they were primarily linguists) which explains how the Germanic languages differ from other Indo-European groups.  For instance, the English word foot is actually the same word as the Latin equivalent, whose stem is ped (as in pedal).  Under Grimm's Law, p mutates into the related sound f, d into the related sound t, and the vowel just mutates.

It goes further.  During the 1st millennium AD, the Western Germanic dialects divided into High German and Low German (that's not like High Elvish, by the way, just a matter of whether they were spoken on the upper or lower Rhine) and foot in Low German (which includes English and Dutch) became fuss (the vowel's pronounced much the same) in High German, which is modern German.  Similarly, in many cases d again turned into t — door/tur, deer/tier etc.

These might seem small changes but, with large numbers of them going on over thousands of years, together with word replacement, languages that started almost the same can change beyond recognition.  In a context without a strong, centralised political or cultural structure, this will take a form where the language spoken in village A and village B diverge a little, although not too much for them to understand one another.  Similarly, the villagers in B won't have too much trouble understanding the speech in village C, but those in A might struggle a bit.  By the time you reach village Z, there appears very little similarity at all.

If, on the other hand, there's a strong state that needs to issue laws and proclamations that will be understood, or if authors and poets are writing works that are understood to be expressions of the entire culture, standard forms will gradually emerge that become national languages.  These often differ little from the language next door, but their speakers like to think of them as different for political or religious reasons: Dutch/Flemish, Serbian/Croatian, Hindi/Urdu, Malay/Indonesian etc.  Curiously, with English and American, which are almost as distinct as some of those pairs, the political separation seems to have fostered a sense of being a common language, instead.

Many of the world's languages are grouped together into widespread super-families — examples include Indo-European, Afro-Asiatic, Altaic, Sino-Tibetan, Austronesian and Niger-Congo, along with many smaller families.  In the New World, the old picture of a dozen or so families has been brought down (though with disagreement from many linguists) to three, with the Amerind family covering all of South and Central America, much of the contiguous US and eastern and central Canada.

The temptation is always to try to link up language families into fewer and larger super-families, but that's not always easy.  Indo-European is a relatively straightforward case, and even that produces controversy and disagreement. 

It's a special case for two reasons.  One is that it's a relatively young family.  Although there are different models for its origin, it's probable that it descends from a group of mutually comprehensible dialects spoken between four and five thousand years ago.  The other is that it includes languages (such as Greek, Latin and Sanskrit) that were extensively written more than half that time ago, and some, such as Hittite, that have left written records from when the family was fairly young.  This makes it quite easy to establish what it was that all these languages are descended from and so identify other members of the family.

On the other hand, many language families are significantly older, and their languages have sometimes only been written down for the first time within the last few centuries.  This makes identifying their relationships even more difficult, especially when they're isolated remnants of families that have largely been replaced by a later wave, a process that's still going on, especially in areas of the world that were colonised by Europeans.  Isolated languages like Basque in the Pyrenees, Burushaski in Kashmir, or the so-called Paleo-Siberian languages may not have had any interaction with any "relatives" that might survive for ten thousand years or more.

Some languages are more conservative than others (Lithuanian, for instance, is often taken as the closest we can get to how the original Indo-European language might have been) but ultimately they all change, mutating their sounds and replacing their vocabulary with foreign imports.  There's a point beyond which current techniques just aren't good enough to detect whether two languages are related or not.

Besides, associations can be as important as "descent", as we saw with the huge quantity of imported vocabulary in English.  Sometimes these influences are greater, affecting even the structure and the stable words used by linguists.  For a long time, there was a controversy over whether Vietnamese was a Kadai language, like its neighbour Lao, or an Austroasiatic language, like its neighbour Khmer.  Most linguists have plumped for the latter, but the language is very much a hybrid.

Japanese is an even stranger case.  The jury's still out here over whether it's an Altaic language that arrived via Korea or an Austronesian language that arrived via the Philippines.  It displays elements of both.

So it's unlikely that we'll ever know for sure whether all human languages ultimately derive from the same source, or if they arose independently in many parts of the world.  The human brain appears to be hard-wired for language, so the multiple invention theory is quite plausible.  It's likely to come down to when our ancestors started talking.  Until quite recently (the idea appears in Pullman's His Dark Materials trilogy) it was believed that language started with a quantum leap in human intellect as recently as 33,000 years ago, at a time when the species was already spread all over the planet, and that would have favoured the multiple origin of language.

Recently, though, both circumstantial evidence of earlier symbolic thought and physical evidence of a speech-enhancing mutation of the larynx have suggested that humans were using language a great deal earlier, probably at a time when they were confined to a relatively small area of Africa.  Although this doesn't prove the single-origin theory, it makes it much more likely.

Then again, maybe they were all taught by Quenya-speaking elves.  It's possible.

So what bearing does any of this have on writing fantasy?  Unless you're going to actually create a whole raft of languages, and then treat your readers to a lecture about them, does it really matter what families they belong to?

Well, yes and no.  It's very much a background aspect, which the reader's unlikely to notice (though it may be far more noticeable if it's ignored or poorly done) but it can contribute to that seamless gloss of reality that the best fantasy worlds somehow achieve.  World-making always reminds me of a swan: seeing it gliding gracefully and effortlessly across the water gives no idea at all of how its legs are going nineteen to the dozen underwater to achieve that impression.

The sounds and elements that make up names can give away subtle relationships between their countries' languages.  Suppose, for instance, several towns in your main country have names ending in -ket — perhaps this is the equivalent of the English element -ton.  Another country (three or four across, perhaps) has a town whose name ends in -gad.  This could suggest the same kind of sound-shift as those described by Grimm's Law, indicating the two countries have related languages, but not closely related.

It can also affect the difficulty characters encounter when learning to speak foreign languages.  It tends to be easier to learn a language if we can latch onto familiar elements and harder if nothing's alike.  I have a scene in an unpublished novel where a character who's good with languages is trying to learn to speak to the people he's staying among.  He comments that it's easier to learn than many, observing that Some of the words seem a bit like Kimdyran.  Like, they say duvin for a cow, and we say tovien.  On the other hand, he tries and totally fails later to learn another, very strange language.  Besides giving an element of his character, it helps to define the relationships between languages, and therefore cultures, within his world.

On the principle that "messy worlds rule", this doesn't have to be too predictable.  In our world, even before languages like English, Spanish and French went global, some language families were extremely far flung — Austronesian, for instance, is spoken all the way from Madagascar to Easter Island, and is mixed up in places with other families.  On the other hand, bordering countries may be linguistically unconnected, such as Hungary, forming a Uralic island in a sea of Indo-European Slavonic languages.

You can perfectly well ignore all this.  Most authors do, I suspect, and often still manage to produce worlds with believable names and cultures.  It can give an extra layer of reality, though, to think about how your languages relate to one another.  And, if you're like me, it's fun.  Which is the main thing, of course.