{"product_id":"formalizing-natural-languages-the-nooj-approach-hardback-9781848219021","title":"Formalizing Natural Languages; The NooJ Approach (Hardback) 9781848219021","description":"\u003cfont face=\"Georgia\"\u003e\r\n\u003cp\u003e\u003cfont size=\"6\"\u003eFormalizing Natural Languages\u003c\/font\u003e\u003cbr\u003e\r\n\u003cfont size=\"5\"\u003eThe NooJ Approach\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\r\n\r\n\r\n\u003cp\u003e\u003cfont size=\"4\"\u003eMax Silberztein (Author)\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\u003cp\u003e\u003cfont size=\"3\"\u003e9781848219021, Wiley\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\u003cp\u003e\u003cfont size=\"3\"\u003eHardback, published 8 January 2016\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\u003cp\u003e\u003cfont size=\"3\"\u003e346 pages\u003cbr\u003e24.1 x 16.5 x 2.3 cm, 0.653 kg\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\r\n\r\n\u003cp align=\"justify\"\u003e\u003cem\u003e\u003cfont size=\"3\"\u003e\u003cp\u003eThis book lays ground for better understanding of both computational linguistics (CL) and natural language processing (NLP) perspectives, i.e. it shows how to describe language (CL) in order to build the best NLP applications (NLP). The book bridges the gap between theoretical linguistic phenomena and practical language models. It shows how computational linguists and language engineers working together can bring us closer to better language understanding by both humans and computers.\u003c\/p\u003e \u003cp\u003eThe author takes us on a stroll through the layers of language processing, explaining very soundly and giving examples and counterexamples that bring additional clarification for each step we make on that path. Starting with the tiny bits of written language, the alphabet, via dictionary and atomic linguistic units that occupy it, he clarifies the importance of each step, giving us solid ground to build upon any language project we might venture to undertake.\u003c\/p\u003e \u003cp\u003eSilberztein knows how to invite an audience into his Project, as he calls it, and introduces the topic in such a manner that makes you want to read the book until the last page (and solve all the CL and NLP problems on the way). He smoothly transitions through Parts one, two and three, building one topic upon the previous one, as if playing with lego blocks.\u003c\/p\u003e \u003cp\u003eHe begins by demonstrating the importance of defining basic (atomic) linguistic units starting with the alphabet and vocabulary that prepare us for the construction of electronic dictionaries. It is the design of the e-dictionary that will allow us and support us in formalizing the language of our interest. Thus, it is not a surprise that a thorough classification and understanding of our basic resources is needed to prepare (and prepare well) and specify affixes [\u003ci\u003ere-, de-, un-, -ation], simple words [home, love, sky], multiword units [sweet potatoes, more and more, round table] and expressions [to give up, to turn off, to take off] that we will play around with to construct and annotate new words, phrases and sentences.\u003c\/i\u003e\u003c\/p\u003e \u003cp\u003eHe then takes regular grammars, context-free grammars, context-sensitive grammars and unrestricted grammars and he makes them all work via NooJ’s multifaceted approach. The (beautiful) simplicity of this application is aligned with the way we, as humans, process vocabulary, grammar, orthography, syntax, semantics…thus making the NooJ as a tool easy to use by beginners and more advanced users alike.\u003c\/p\u003e \u003cp\u003eIt is only expected that the journey will end with applications both in parsing and generating written text. We are presented with the lexical analysis, syntactic analysis (local and structural) and transformational analysis that open up the door for more sophisticated NLP applications (Question Answering, Machine Translation, Semantic Analyzer, etc.)\u003c\/p\u003e \u003cp\u003eThe most expected audience of ‘\"Formalizing Natural Languages: The NooJ Approach’ are linguists i.e. computational linguists and NLP people (or as the author likes to call them language engieers). But, since the book holds the key that can open a whole sea of possible applications in the domains of other subfields, I would recommend it to etymologists, sociolinguists, psycholinguists, forensic linguists, internet linguists, corpus linguists or to any data scientist today. Having each chapter end with exercises and additional internet links, the book is also suitable as a class reading in NLP and CL classes, machine translation and similar. The book is presented in a way as to improve the understanding of the ways the natural language can be formalized and has the power to reveal some new applications to almost any type of written text. Since the book and NooJ as a tool came into existence in the era dominated by unstructured data, the potential of presented tool is limited only by the imagination of its user. \u003cbr\u003e—Kristina Kocijan, Department of Information and Communication Sciences Faculty of Humanities and Social Sciences University of Zagreb, Croatia\u003c\/p\u003e\u003c\/font\u003e\u003c\/em\u003e\u003c\/p\u003e\r\n\r\n\u003cp align=\"justify\"\u003e\u003cstrong\u003e\u003cfont size=\"3\"\u003e\u003cp\u003eThis book is at the very heart of linguistics. It provides the theoretical and methodological framework needed to create a successful linguistic project.\u003c\/p\u003e \u003cp\u003ePotential applications of descriptive linguistics include spell-checkers, intelligent search engines, information extractors and annotators, automatic summary producers, automatic translators, and more. These applications have considerable economic potential, and it is therefore important for linguists to make use of these technologies and to be able to contribute to them.\u003c\/p\u003e \u003cp\u003eThe author provides linguists with tools to help them formalize natural languages and aid in the building of software able to automatically process texts written in natural language (Natural Language Processing, or NLP).\u003c\/p\u003e \u003cp\u003eComputers are a vital tool for this, as characterizing a phenomenon using mathematical rules leads to its formalization. NooJ – a linguistic development environment software developed by the author – is described and practically applied to examples of NLP.\u003c\/p\u003e\u003c\/font\u003e\u003c\/strong\u003e\u003c\/p\u003e\r\n\r\n\u003cp\u003e\u003cfont size=\"3\"\u003e\u003cp\u003eAcknowledgments xi\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 1. Introduction: the Project 1\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e1.1. Characterizing a set of infinite size 4\u003c\/p\u003e \u003cp\u003e1.2. Computers and linguistics 5\u003c\/p\u003e \u003cp\u003e1.3. Levels of formalization 6\u003c\/p\u003e \u003cp\u003e1.4. Not applicable 7\u003c\/p\u003e \u003cp\u003e1.4.1. Poetry and plays on words 7\u003c\/p\u003e \u003cp\u003e1.4.2. Stylistics and rhetoric 9\u003c\/p\u003e \u003cp\u003e1.4.3. Anaphora, coreference resolution, and semantic disambiguation 10\u003c\/p\u003e \u003cp\u003e1.4.4. Extralinguistic calculations 12\u003c\/p\u003e \u003cp\u003e1.5. NLP applications 12\u003c\/p\u003e \u003cp\u003e1.5.1. Automatic translation 14\u003c\/p\u003e \u003cp\u003e1.5.2. Part-of-speech (POS) tagging 18\u003c\/p\u003e \u003cp\u003e1.5.3. Linguistic rather than stochastic analysis 27\u003c\/p\u003e \u003cp\u003e1.6. Linguistic formalisms: NooJ 27\u003c\/p\u003e \u003cp\u003e1.7. Conclusion and structure of this book 30\u003c\/p\u003e \u003cp\u003e1.8. Exercises 31\u003c\/p\u003e \u003cp\u003e1.9. Internet links 32\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart 1. Linguistic Units 35\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 2. Formalizing the Alphabet 37\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e2.1. Bits and bytes 37\u003c\/p\u003e \u003cp\u003e2.2. Digitizing information 39\u003c\/p\u003e \u003cp\u003e2.3. Representing natural numbers 39\u003c\/p\u003e \u003cp\u003e2.3.1. Decimal notation 39\u003c\/p\u003e \u003cp\u003e2.3.2. Binary notation 40\u003c\/p\u003e \u003cp\u003e2.3.3. Hexadecimal notation 41\u003c\/p\u003e \u003cp\u003e2.4. Encoding characters 41\u003c\/p\u003e \u003cp\u003e2.4.1. Standardization of encodings 43\u003c\/p\u003e \u003cp\u003e2.4.2. Accented Latin letters, diacritical marks, and ligatures 45\u003c\/p\u003e \u003cp\u003e2.4.3. Extended ASCII encodings 46\u003c\/p\u003e \u003cp\u003e2.4.4. Unicode 47\u003c\/p\u003e \u003cp\u003e2.5. Alphabetical order 53\u003c\/p\u003e \u003cp\u003e2.6. Classification of characters 56\u003c\/p\u003e \u003cp\u003e2.7. Conclusion 56\u003c\/p\u003e \u003cp\u003e2.8. Exercises 57\u003c\/p\u003e \u003cp\u003e2.9. Internet links 57\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 3. Defining Vocabulary 59\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e3.1. Multiple vocabularies and the evolution of vocabulary 59\u003c\/p\u003e \u003cp\u003e3.2. Derivation 63\u003c\/p\u003e \u003cp\u003e3.2.1. Derivation applies to vocabulary elements 63\u003c\/p\u003e \u003cp\u003e3.2.2. Derivations are unpredictable 64\u003c\/p\u003e \u003cp\u003e3.2.3. Atomicity of derived words 65\u003c\/p\u003e \u003cp\u003e3.3. Atomic linguistic units (ALUs) 67\u003c\/p\u003e \u003cp\u003e3.3.1. Classification of ALUs 67\u003c\/p\u003e \u003cp\u003e3.4. Multiword units versus analyzable sequences of simple words 70\u003c\/p\u003e \u003cp\u003e3.4.1. Semantics 72\u003c\/p\u003e \u003cp\u003e3.4.2. Usage 76\u003c\/p\u003e \u003cp\u003e3.4.3. Transformational analysis 77\u003c\/p\u003e \u003cp\u003e3.5. Conclusion 80\u003c\/p\u003e \u003cp\u003e3.6. Exercises 81\u003c\/p\u003e \u003cp\u003e3.7. Internet links 81\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 4. Electronic Dictionaries 83\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e4.1. Could editorial dictionaries be reused? 83\u003c\/p\u003e \u003cp\u003e4.2. LADL electronic dictionaries 90\u003c\/p\u003e \u003cp\u003e4.2.1. Lexicon-grammar 90\u003c\/p\u003e \u003cp\u003e4.2.2. DELA 93\u003c\/p\u003e \u003cp\u003e4.3. Dubois and Dubois-Charlier electronic dictionaries 94\u003c\/p\u003e \u003cp\u003e4.3.1. The Dictionnaire électronique des mots 95\u003c\/p\u003e \u003cp\u003e4.3.2. Les Verbes Français (LVF) 97\u003c\/p\u003e \u003cp\u003e4.4. Specifications for the construction of an electronic dictionary 99\u003c\/p\u003e \u003cp\u003e4.4.1. One ALU = one lexical entry 99\u003c\/p\u003e \u003cp\u003e4.4.2. Importance of derivation 100\u003c\/p\u003e \u003cp\u003e4.4.3. Orthographic variation 101\u003c\/p\u003e \u003cp\u003e4.4.4. Inflection of simple words, compound words, and expressions 103\u003c\/p\u003e \u003cp\u003e4.4.5. Expressions 104\u003c\/p\u003e \u003cp\u003e4.4.6. Integration of syntax and semantics 104\u003c\/p\u003e \u003cp\u003e4.5. Conclusion 107\u003c\/p\u003e \u003cp\u003e4.6. Exercises  108\u003c\/p\u003e \u003cp\u003e4.7. Internet links 108\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart 2. Languages, Grammars and Machines 111\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 5. Languages, Grammars, and Machines  113\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e5.1. Definitions 113\u003c\/p\u003e \u003cp\u003e5.1.1. Letters and alphabets 113\u003c\/p\u003e \u003cp\u003e5.1.2. Words and languages 114\u003c\/p\u003e \u003cp\u003e5.1.3. ALU, vocabularies, phrases, and languages 114\u003c\/p\u003e \u003cp\u003e5.1.4. Empty string 115\u003c\/p\u003e \u003cp\u003e5.1.5. Free language 116\u003c\/p\u003e \u003cp\u003e5.1.6. Grammars 116\u003c\/p\u003e \u003cp\u003e5.1.7. Machines 117\u003c\/p\u003e \u003cp\u003e5.2. Generative grammars 118\u003c\/p\u003e \u003cp\u003e5.3. Chomsky-Schützenberger hierarchy 119\u003c\/p\u003e \u003cp\u003e5.3.1. Linguistic formalisms 122\u003c\/p\u003e \u003cp\u003e5.4. The NooJ approach  124\u003c\/p\u003e \u003cp\u003e5.4.1. A multifaceted approach 124\u003c\/p\u003e \u003cp\u003e5.4.2. Unified notation 125\u003c\/p\u003e \u003cp\u003e5.4.3. Cascading architecture 127\u003c\/p\u003e \u003cp\u003e5.5. Conclusion 127\u003c\/p\u003e \u003cp\u003e5.6. Exercises  128\u003c\/p\u003e \u003cp\u003e5.7. Internet links 129\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 6. Regular Grammars 131\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e6.1. Regular expressions 131\u003c\/p\u003e \u003cp\u003e6.1.1. Some examples of regular expressions 135\u003c\/p\u003e \u003cp\u003e6.2. Finite-state graphs 137\u003c\/p\u003e \u003cp\u003e6.3. Non-deterministic and deterministic graphs 139\u003c\/p\u003e \u003cp\u003e6.4. Minimal deterministic graphs 141\u003c\/p\u003e \u003cp\u003e6.5. Kleene’s theorem 142\u003c\/p\u003e \u003cp\u003e6.6. Regular expressions with outputs and finite-state transducers 146\u003c\/p\u003e \u003cp\u003e6.7. Extensions of regular grammars 151\u003c\/p\u003e \u003cp\u003e6.7.1. Lexical symbols 151\u003c\/p\u003e \u003cp\u003e6.7.2. Syntactic symbols 153\u003c\/p\u003e \u003cp\u003e6.7.3. Symbols defined by grammars 154\u003c\/p\u003e \u003cp\u003e6.7.4. Special operators 155\u003c\/p\u003e \u003cp\u003e6.8. Conclusion 159\u003c\/p\u003e \u003cp\u003e6.9. Exercises 159\u003c\/p\u003e \u003cp\u003e6.10. Internet links 159\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 7. Context-Free Grammars 161\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e7.1. Recursion 164\u003c\/p\u003e \u003cp\u003e7.1.1. Right recursion 166\u003c\/p\u003e \u003cp\u003e7.1.2. Left recursion 167\u003c\/p\u003e \u003cp\u003e7.1.3. Middle recursion 168\u003c\/p\u003e \u003cp\u003e7.2. Parse trees  170\u003c\/p\u003e \u003cp\u003e7.3. Conclusion 173\u003c\/p\u003e \u003cp\u003e7.4. Exercises 173\u003c\/p\u003e \u003cp\u003e7.5. Internet links 174\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 8. Context-Sensitive Grammars 175\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e8.1. The NooJ approach  176\u003c\/p\u003e \u003cp\u003e8.1.1. The anbncn language 177\u003c\/p\u003e \u003cp\u003e8.1.2. The language a2n 180\u003c\/p\u003e \u003cp\u003e8.1.3. Handling reduplications 181\u003c\/p\u003e \u003cp\u003e8.1.4. Grammatical agreements 182\u003c\/p\u003e \u003cp\u003e8.1.5. Lexical constraints in morphological grammars 185\u003c\/p\u003e \u003cp\u003e8.2. NooJ contextual constraints 186\u003c\/p\u003e \u003cp\u003e8.3. NooJ variables 188\u003c\/p\u003e \u003cp\u003e8.3.1. Variables’ scope 188\u003c\/p\u003e \u003cp\u003e8.3.2. Computing a variable’s value 189\u003c\/p\u003e \u003cp\u003e8.3.3. Inheriting a variable’s value 191\u003c\/p\u003e \u003cp\u003e8.4. Conclusion 191\u003c\/p\u003e \u003cp\u003e8.5. Exercises  192\u003c\/p\u003e \u003cp\u003e8.6. Internet links 192\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 9. Unrestricted Grammars 195\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e9.1. Linguistic adequacy 197\u003c\/p\u003e \u003cp\u003e9.2. Conclusion 199\u003c\/p\u003e \u003cp\u003e9.3. Exercise 199\u003c\/p\u003e \u003cp\u003e9.4. Internet links 199\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart 3. Automatic Linguistic Parsing 201\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 10. Text Annotation Structure 205\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e10.1. Parsing a text 205\u003c\/p\u003e \u003cp\u003e10.2. Annotations 206\u003c\/p\u003e \u003cp\u003e10.2.1. Limits of XML\/TEI representation 207\u003c\/p\u003e \u003cp\u003e10.3. Text annotation structure (TAS) 208\u003c\/p\u003e \u003cp\u003e10.4. Exercise 211\u003c\/p\u003e \u003cp\u003e10.5. Internet links 212\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 11. Lexical Analysis 213\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e11.1. Tokenization 213\u003c\/p\u003e \u003cp\u003e11.1.1. Letter recognition 214\u003c\/p\u003e \u003cp\u003e11.1.2. Apostrophe\/quote 217\u003c\/p\u003e \u003cp\u003e11.1.3. Dash\/hyphen 219\u003c\/p\u003e \u003cp\u003e11.1.4. Dot\/period\/point ambiguity 222\u003c\/p\u003e \u003cp\u003e11.2. Word forms  224\u003c\/p\u003e \u003cp\u003e11.2.1. Space and punctuation 224\u003c\/p\u003e \u003cp\u003e11.2.2. Numbers 226\u003c\/p\u003e \u003cp\u003e11.2.3. Words in upper case 228\u003c\/p\u003e \u003cp\u003e11.3. Morphological analyses 229\u003c\/p\u003e \u003cp\u003e11.3.1. Inflectional morphology  230\u003c\/p\u003e \u003cp\u003e11.3.2. Derivational morphology 234\u003c\/p\u003e \u003cp\u003e11.3.3. Lexical morphology  236\u003c\/p\u003e \u003cp\u003e11.3.4. Agglutinations  239\u003c\/p\u003e \u003cp\u003e11.4. Multiword unit recognition 241\u003c\/p\u003e \u003cp\u003e11.5. Recognizing expressions 243\u003c\/p\u003e \u003cp\u003e11.5.1. Characteristic constituent 244\u003c\/p\u003e \u003cp\u003e11.5.2. Varying the characteristic constituent 245\u003c\/p\u003e \u003cp\u003e11.5.3. Varying the light verb 246\u003c\/p\u003e \u003cp\u003e11.5.4. Resolving ambiguity 247\u003c\/p\u003e \u003cp\u003e11.5.5. Annotating expressions 251\u003c\/p\u003e \u003cp\u003e11.6. Conclusion  254\u003c\/p\u003e \u003cp\u003e11.7. Exercise  255\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 12. Syntactic Analysis 257\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e12.1. Local grammars 257\u003c\/p\u003e \u003cp\u003e12.1.1. Named entities 257\u003c\/p\u003e \u003cp\u003e12.1.2. Grammatical word sequences 262\u003c\/p\u003e \u003cp\u003e12.1.3. Automatically identifying ambiguity 263\u003c\/p\u003e \u003cp\u003e12.2. Structural grammars 265\u003c\/p\u003e \u003cp\u003e12.2.1. Complex atomic linguistic units 266\u003c\/p\u003e \u003cp\u003e12.2.2. Structured annotations 268\u003c\/p\u003e \u003cp\u003e12.2.3. Ambiguities 270\u003c\/p\u003e \u003cp\u003e12.2.4. Syntax trees vs parse trees 273\u003c\/p\u003e \u003cp\u003e12.2.5. Dependency grammar and tree 276\u003c\/p\u003e \u003cp\u003e12.2.6. Resolving ambiguity transparently 279\u003c\/p\u003e \u003cp\u003e12.3. Conclusion 280\u003c\/p\u003e \u003cp\u003e12.4. Exercises 281\u003c\/p\u003e \u003cp\u003e12.5. Internet links 281\u003c\/p\u003e \u003cp\u003e\u003cb\u003eChapter 13. Transformational Analysis 283\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e13.1. Implementing transformations 286\u003c\/p\u003e \u003cp\u003e13.2. Theoretical problems 292\u003c\/p\u003e \u003cp\u003e13.2.1. Equivalence of transformation sequences 292\u003c\/p\u003e \u003cp\u003e13.2.2. Ambiguities in transformed sentences 293\u003c\/p\u003e \u003cp\u003e13.2.3. Theoretical sentences 294\u003c\/p\u003e \u003cp\u003e13.2.4. The number of transformations to be implemented 295\u003c\/p\u003e \u003cp\u003e13.3. Transformational analysis with NooJ 297\u003c\/p\u003e \u003cp\u003e13.3.1. Applying a grammar in “generation” mode 298\u003c\/p\u003e \u003cp\u003e13.3.2. The transformation’s arguments 299\u003c\/p\u003e \u003cp\u003e13.4. Question answering 303\u003c\/p\u003e \u003cp\u003e13.5. Semantic analysis 304\u003c\/p\u003e \u003cp\u003e13.6. Machine translation 305\u003c\/p\u003e \u003cp\u003e13.7. Conclusion 309\u003c\/p\u003e \u003cp\u003e13.8. Exercises 309\u003c\/p\u003e \u003cp\u003e13.9. Internet links 310\u003c\/p\u003e \u003cp\u003eConclusion 311\u003c\/p\u003e \u003cp\u003eBibliography 315\u003c\/p\u003e \u003cp\u003eIndex 327\u003c\/p\u003e\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\u003cp\u003e\u003cfont size=\"3\"\u003eSubject Areas: Linguistics [\u003ca title=\"See our other books on Linguistics\" href=\"https:\/\/freshlyprintedbooks.co.uk\/search?q=%22Linguistics%20%5BCF%5D%22\"\u003eCF\u003c\/a\u003e]\u003c\/font\u003e\u003c\/p\u003e\r\n\r\n\r\n\u003c\/font\u003e","brand":"Wiley-ISTE","offers":[{"title":"Brand New","offer_id":52449400226072,"sku":"9781848219021","price":122.99,"currency_code":"GBP","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0730\/2037\/5320\/files\/9781848219021.jpg?v=1785198128","url":"https:\/\/freshlyprintedbooks.co.uk\/products\/formalizing-natural-languages-the-nooj-approach-hardback-9781848219021","provider":"Freshly Printed Books","version":"1.0","type":"link"}