[Lingtyp] IPA for linguists' names

Yuri Koryakov ybkoryakov at gmail.com
Mon Sep 7 17:00:25 UTC 2026


Dear Mattis,

On 07.09.2026 10:46, Mattis List via Lingtyp wrote:
> Two points here. First, if you want to archive this project, Yuri, I 
> suggest to upload the data after collection to a public repository. I 
> suggest Zenodo.org, and share the DOI here, but please also add a 
> clear license that would guarantee the re-use of the data. Licensing 
> is extremely important to make sure people could potentially remix 
> this in future projects that might take up the idea by Martin and 
> provide IPA for languages as well.
Yes, thanks for the advice, I'll probably do that! cc-by-sa will 
probably suffice.
>
> Second, using Google Sheets is also excluding scholars from China. An 
> open server solution would be better. I mention this, since it has not 
> been brought up so far. But if the idea is anyway to collect data and 
> then make a dump, this should suffice for now.

For scholars from China, it might be worth creating an additional place 
to collect data on another resource, similar to Google Spreadsheets. But 
this requires help from people from China.
>
> Anyway: the archiving with a repository seems important, to guarantee 
> other people may build on it.
>
> Best,
>
> Mattis
Yours,
Yuri
>
> On 06.09.26 21:18, Yuri Koryakov via Lingtyp wrote:
>> Dear Jocelyn,
>>
>> On 06.09.2026 13:39, Jocelyn Aznar via Lingtyp wrote:
>>> Dear Yuri,
>>>
>>> Le 06/09/2026 à 07:25, Yuri Koryakov via Lingtyp a écrit :
>>>> On the one hand, all information not added by people themselves is 
>>>> taken from public online sources; people add further information 
>>>> voluntarily. On the other hand, this table does not contain any 
>>>> sensitive personal information – no emails, place of residence, 
>>>> place of work, position, etc.
>>>
>>> Note that you suggest to people to provide email address in the 
>>> comment section.
>>
>> Yes, you're right. Although this only applies to quite rare cases 
>> where there are complete namesakes and they need to be distinguished 
>> somehow, and only as one of the methods. I've added ORCID, by the 
>> way, thanks for the idea!
>>
>>> In some countries (like the US), the pronoun to use is actually a 
>>> sensitive personal information, as it reveal gender info. Native 
>>> language can be a proxy for ethnic identity which also can be 
>>> problematic in many countries. Google Doc, especially with a free 
>>> account, is not safe for personal info, not to mention that 
>>> someone/AI could vandalize the spreadsheet and insert/modify the 
>>> data, and when collecting personal info, the database creator is 
>>> responsible of providing a database which is secured enough 
>>> according to the sensitivity of the info collected.
>>
>> Yes, it's very important. For a while, I'll keep an eye out for 
>> vandalism by checking the edit history.  After the active phase of 
>> adding names, I will make a static webpage and change the way new 
>> names are added.
>>
>>> More practical issue, I can imagine in the US that inviting a 
>>> researcher that asked for the pronoun they could be problematic, 
>>> either for the researcher or the lab inviting them. As it is 
>>> information available on Google Doc, this will be automatically 
>>> scanned by the IA that the border police uses.
>>
>> Horrible, I didn't know about such problems in the US (I would 
>> understand if it were Russia...).
>>
>>> Birthday can also be problematic actually, as the other 
>>> disambiguators, as they allow to identify the person, which is 
>>> normal for the purpose of this database, two people having the same 
>>> name might want different ways of being addressed. Better in this 
>>> case ORCID or professional URL which are publicly available data.
>>>
>>> Somehow, GDPR might seem annoying, but, in this day and age of IA, 
>>> any info available on internet is being scrapped and compiled, and 
>>> automatically analyzed by more and more fascist governments and 
>>> organizations for whatever reasons they have, or will have. Somehow, 
>>> I wonder more and more about the relevance of open access data in 
>>> this new context.
>>
>> To sum up, I can say that all sensitive information is added by 
>> people themselves. Just in case, I've added the following warning: 
>> *Warning*: All information you provide will be publicly available. If 
>> you feel any information is too sensitive and/or could be used 
>> against you, do not provide it.
>>
>>>>> - There isn't only one valid pronunciation of my name.
>>>>
>>>> The IPA column allows you to specify multiple pronunciation 
>>>> variants, both equivalent and language-specific.
>>>
>>> Right.
>>>
>>>>>
>>>>> - I'm ok with people having difficulties pronouncing my name (at 
>>>>> the end of the day, no one can master the pronunciation of the 
>>>>> many phonologies we work with in our profession).
>>>>> - Pushing people to pronounce your name following your 
>>>>> expectations can cause situations in which they feel really ill at 
>>>>> ease because they cannot pronounce your name.
>>>> Well, we don't force people to pronounce the names exactly this 
>>>> way, but we do give them a reference point. For example, when I 
>>>> just see the name Jocelyn, I don't even know what the first sound 
>>>> might be – [ʤ / ʒ / x / j] or something else.
>>>
>>> Of course I understand this issue.
>>>
>>>>> - I'm also ok with people adapting my name if it is too hard to 
>>>>> pronounce, that's usually what happens in practice.
>>>>> - From a more technical angle,it would make more sense if the IPA 
>>>>> transcription follows the same order of the first name + other name.
>>>> Okay, maybe. But this way one can show the order in which a person 
>>>> prefers to pronounce their name. (By the way, the most common order 
>>>> in this column is: first name + other name.)
>>>
>>> This is not so clear from the instruction neither, and this is not 
>>> what's systematically done from the data that was already populated.
>> The Guide says "Please provide a transcription of your full name (*in 
>> your preferred order*)."
>>>
>>> For the purpose of this database, I would have a dedicated column 
>>> for the preferred name for a seminar/event, not just "Other names" 
>>> which can lead to confusion as it might contain many info in 
>>> different order. 1 info : 1 column.
>> If I created a separate column for each type of information, the 
>> database would grow too large. And in general, a person can provide 
>> the preferred name for a seminar/event when registering for that 
>> seminar/event.
>>> I think this project would be better handled through the server of 
>>> an institution (well, more of our institutions rely on 
>>> Microsoft/Google and co., but if they pay for the services, the data 
>>> might be actually properly handled), with a better interface, like a 
>>> form in which people can enter these info and a page that displays 
>>> those info.
>> Yes, perhaps that would be the ideal solution. But I don't have such 
>> possibility (and I'm not sure all linguists would be willing to 
>> provide such information on a Russian institution's server).  I 
>> proceed from a simple maxim: The best is the enemy of the good.
>>>
>>> Best,
>>> Jocelyn
>> Yours,
>> Yuri
>>> _______________________________________________
>>> Lingtyp mailing list
>>> Lingtyp at listserv.linguistlist.org
>>> https://listserv.linguistlist.org/cgi-bin/mailman/listinfo/lingtyp
>>
>> _______________________________________________
>> Lingtyp mailing list
>> Lingtyp at listserv.linguistlist.org
>> https://listserv.linguistlist.org/cgi-bin/mailman/listinfo/lingtyp
>
> _______________________________________________
> Lingtyp mailing list
> Lingtyp at listserv.linguistlist.org
> https://listserv.linguistlist.org/cgi-bin/mailman/listinfo/lingtyp


More information about the Lingtyp mailing list