7 ms·
> Unicode 13.0 adds 5,930 characters, for a total of 143,859 characters. These additions include 4 new scripts, for a total of 154 scripts, as well as 55 new em
by russellallen 7y ago
> Unicode 13.0 adds 5,930 characters, for a total of 143,859 characters. These additions include 4 new scripts, for a total of 154 scripts, as well as 55 new emoji characters.
So how far off is Unicode from being 'done'? At what point will they be able to stop adding characters and scripts?
- tambre 7y agoOnce people stop needing and/or inventing new characters and scripts.
- magicalhippo 7y agoAnd emojis, don't forget about the all-important emojis.
- samastur 7y agoYou are probably joking, but I am seriously annoyed there's no donkey emoji.
- rbonvall 7y agoMy wife and I call each other "donkey", and some years ago we used the horse emoji, which was low resolution enough to look as a donkey if you squinted. But modern emojis are too high resolution and it definitely looks like a horse now. So I feel your pain.
- JdeBP 7y agoIs U+130D8 not good enough for you?
- mrspuratic 7y agoIf there's room in Unicode to describe "ARABIC LETTER BEH WITH THREE DOTS POINTING UPWARDS BELOW AND TWO DOTS ABOVE" there's room for "DONKEY". I wonder why Gardiner's descriptions never made it into Unicode? https://en.wikipedia.org/wiki/List_of_Egyptian_hieroglyphs#E https://en.wikipedia.org/wiki/List_of_Egyptian_hieroglyphs#E O, wait, there are censored genitalia in D block ...
- samastur 7y agoThanks, I wasn't aware of it. Sadly I get square when I try to use it so less reliable than emoji.
- mrspuratic 7y agoI wanted a pink pony. My wifi SSID is <horse U+1F40E><unicorn U+1F984>, which is a barely satisfactory approximation. (FWIW the FTP site is still up: ftp://ftp.unicode.org/Public/13.0.0/ )
- AnIdiotOnTheNet 7y agoEmoji: It's like Kanji, only without agreed upon semantic meaning or pronunciation. I pity historians of the future who have to try and decipher this garbage.
- tialaramex 7y agoHa, imagine if some idiots designed a whole country so that its correct functioning depended upon trying to decipher the meaning of symbols which lack semantic meaning decades or even centuries later. Oh wait, that's the United States of America. What does constitute a "well regulated militia" and how exactly did those symbols come to mean people you wouldn't trust with a sharp stick are allowed semi-automatic hand guns? Symbols mean whatever people want them to mean, that's the extent to which Humpty is right when he lectures Alice. Whether it's an Eggplant emoji or "Ugandan discussions" the symbol is not the meaning of the symbol.
- rbanffy 7y ago"214 graphic characters that provide compatibility with various home computers from the mid-1970s to the mid-1980s and with early teletext broadcasting standards" This part is dear to me, as I helped craft it. It includes 2x3 videotext mosaic characters that will make it much easier to draw large text have have better quality charts in text terminal interfaces. And, of course, the ability to properly encode documents that were generated in computers in the 70's and 80's that contained those platform-specific characters. For 14 we are planning on adding symbols from the Sharp MZ series and the large text characters (3x3 cells) of HP terminals.
- bhaak 7y agoThe niche audience of terminal based games will also be thankful forever to you and the rest of the people that got those characters into Unicode.
- rbanffy 7y agoGo get them. https://www.unicode.org/charts/PDF/U1FB00.pdf https://www.unicode.org/charts/PDF/U1FB00.pdf
- phaker 7y agoDown for me as is everything else. Found a unicode consortium tweet (!) with a picture of the whole block: https://twitter.com/unicode/status/1085613123183071232?lang=en https://twitter.com/unicode/status/1085613123183071232?lang=...
- JdeBP 7y agoThey might be interested in Unscii, too. It has been updated in light of Unicode 13. * http://pelulamu.net/unscii/ http://pelulamu.net/unscii/ (https://news.ycombinator.com/item?id=18478350 https://news.ycombinator.com/item?id=18478350)
- pmarreck 7y agoDo these... do these include the old Commodore 64 metacharacters/symbols??
- pilif 7y agoeither that or when there are 2^24 used codepoints.
- klodolph 7y agoThe available space is closer to 2^20 (0-10FFFF, minus surrogate pairs, depending on whether you are talking about Unicode scalar values or code points).
- Someone 7y agoThere’s also Emoji modifiers (https://en.wikipedia.org/wiki/Miscellaneous_Symbols_and_Pictographs#Emoji_modifiers https://en.wikipedia.org/wiki/Miscellaneous_Symbols_and_Pict...) and regional indicators https://en.wikipedia.org/wiki/Regional_Indicator_Symbol https://en.wikipedia.org/wiki/Regional_Indicator_Symbol that complicate determining the number of Unicode characters.
- taejo 7y agoGerman, a European language that has been more or less standardized for several centuries, with a Latin-based alphabet, added a new letter (ẞ) to its alphabet in 2017. As long as that continues to happen, Unicode will have to add new characters, even if no more ancient scripts are discovered and no new writing systems are developed for currently unwritten languages.
- tasogare 7y agoSmall precision for those who don't know the context: the Eszett (which comes from the ligature of 'ss') existed for centuries already in German writing. 2017 is just the date of its official integration in the alphabet, so it's not a 'new' letter created from scratch. I say that because I remember learning it at school a few decades ago (even if at the time we were warned the subject was touchy), and I was surprised it wasn't standardized that earlier.
- oprypin 7y agoThis is about capital ẞ. Small ß must've been in Unicode from the start.
- Sukram21 7y agoThe Eszett (ß) has been already standardized for several decades (since 1986: with ISO 8859-1 aka latin-1). The newly added letter is the "capital letter Eszett", which did not exist until recently. One could argue that this new letter is not really needed, as Eszett does not appear in capitalized form except when a word is in all-caps, and was then simply written as "SS".
- chrisseaton 7y ago> So how far off is Unicode from being 'done'? Do you think human written language is ‘done’ and will never evolve?
- ken 7y agoI think that most of the changes in Unicode 13 are not from the evolution of human written language. I don't know anyone who's ever written "blueberries" by drawing a picture of some blueberries in the middle of their text.
- nradov 7y agoEmojis literally are an evolution in human written language. They started with youth texting and are now showing up in business emails. I predict that within 50 years we'll see emojis as a routine component of New York Times articles.
- pmarreck 7y agoI'm not going to hold my breath on this one. Emojis have an air of informality that is not appropriate in many circumstances. Imagine writing a death notice with emojis
- JdeBP 7y agoWell one could use U+1303F. But you wait until Maya script gets into Unicode. You'll have at least three different Maya codepoints for death. (-:
- ken 7y agoNow you're getting into the definition of "writing". I would say I've only seen emoji typed, not written. (Before anyone asks: yes, I've seen cuneiform written. I have some interesting friends.) If you count any visual communication that is typed on a phone under the greater umbrella of "writing", then we could also include colors, styles, orientation, funny fonts, image memes, animation, etc. There's no end to the possible visual communication that people might want to transmit digitally. Where do you draw the line? I draw it at "anything in or using a language that people might write in the absence of computers, which they would then reasonably want to store and transmit using a computer". I don't include "any possible visual communication that can occur using a computer". That's far too broad to define "text", or be part of any existing "language", which are the stated goals of Unicode.
- throw0101a 7y agoUntil there's a kumquat emoji then it will not be done: * https://www.emojis.com/food/fruit/ https://www.emojis.com/food/fruit/ There will always be another Emoji that someone, somewhere wants to add.
- ygra 7y agoWell, it was a can of worms once they took the existing characters which for the most part were culturally very tied to Japan. Now everyone had access to those fun little icons and in turn a lot of people felt concepts from their cultural surrounding underrepresented. And while every Unicode announcement gets derided because they added more emoji, it's still just a small subset of the standard. Back when they added them I wasn't much of a fan, but by now I think it was the right decision. I still don't use them, but I've heard they're quite popular in younger age groups.
- jcranmer 7y agoI believe all scripts that are in extant use today are complete; if not, the missing extant scripts are down to the scripts with very few literate people (thousands or fewer). Many of the additions today, barring emoji, are covering historical usage. This includes things like Medieval scribal annotations, a different set of numbers for the Ottomans, and the Mayan script. It will still be over a decade for the historical work to be complete, since there is often a lot of actual research that needs to be done to understand how an ancient writing system works, which has to come before you can even put together a coherent proposal for a new script.
- NelsonMinar 7y agoI think that's a good framework to think about Unicode being "done". But even extant scripts aren't done; consider the Bopomofo additions here in Unicode 13. And it's not clear what "done" even means for Chinese characters.
- Analemma_ 7y agoThe Script Encoding Initiative [0] is a UC Berkeley project to add Unicode support for uncommon and historical scripts. They have a list of remaining scripts which is encouragingly short [1], and most of the scripts on that list have proposals in progress so in theory they should be adopted soon [0]: https://linguistics.berkeley.edu/sei/index.html https://linguistics.berkeley.edu/sei/index.html [1]: https://linguistics.berkeley.edu/sei/scripts-not-encoded.html https://linguistics.berkeley.edu/sei/scripts-not-encoded.htm...