@polycode-projects/the-mechanical-code-talker 2.3.0 → 2.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/corpus/LICENSES.json +19 -4
- package/corpus/README.md +48 -0
- package/corpus/generated/README.md +24 -9
- package/corpus/generated/ace-surface-variants.jsonl +4 -1
- package/corpus/generated/manifest.json +4 -4
- package/corpus/prose/manifest.json +512 -0
- package/corpus/prose/sqlite/LICENSE-NOTICE +53 -0
- package/corpus/prose/sqlite/arch.txt +213 -0
- package/corpus/prose/sqlite/atomiccommit.txt +1117 -0
- package/corpus/prose/sqlite/faq.txt +473 -0
- package/corpus/prose/sqlite/fileformat.txt +1589 -0
- package/corpus/prose/sqlite/lang_createtable.txt +1339 -0
- package/corpus/prose/sqlite/lang_insert.txt +580 -0
- package/corpus/prose/sqlite/lang_select.txt +3293 -0
- package/corpus/prose/sqlite/optoverview.txt +908 -0
- package/corpus/prose/sqlite/queryplanner.txt +447 -0
- package/corpus/prose/sqlite/transactional.txt +41 -0
- package/corpus/prose/sqlite/wal.txt +567 -0
- package/corpus/prose/sqlite/whentouse.txt +300 -0
- package/corpus/prose/wikipedia/Apple.txt +4 -0
- package/corpus/prose/wikipedia/Attempto_Controlled_English.txt +169 -0
- package/corpus/prose/wikipedia/Automated_planning_and_scheduling.txt +67 -0
- package/corpus/prose/wikipedia/Bee.txt +7 -0
- package/corpus/prose/wikipedia/Bird.txt +8 -0
- package/corpus/prose/wikipedia/Bone.txt +4 -0
- package/corpus/prose/wikipedia/Book.txt +7 -0
- package/corpus/prose/wikipedia/Bread.txt +6 -0
- package/corpus/prose/wikipedia/Butterfly.txt +6 -0
- package/corpus/prose/wikipedia/Car.txt +1 -0
- package/corpus/prose/wikipedia/Cat.txt +1 -0
- package/corpus/prose/wikipedia/Child.txt +3 -0
- package/corpus/prose/wikipedia/City.txt +2 -0
- package/corpus/prose/wikipedia/Clock.txt +2 -0
- package/corpus/prose/wikipedia/Cooking.txt +1 -0
- package/corpus/prose/wikipedia/Description_logic.txt +660 -0
- package/corpus/prose/wikipedia/Doctor.txt +6 -0
- package/corpus/prose/wikipedia/Dog.txt +4 -0
- package/corpus/prose/wikipedia/Eagle.txt +4 -0
- package/corpus/prose/wikipedia/Emotion.txt +9 -0
- package/corpus/prose/wikipedia/Eye.txt +5 -0
- package/corpus/prose/wikipedia/Family.txt +3 -0
- package/corpus/prose/wikipedia/Farm.txt +4 -0
- package/corpus/prose/wikipedia/Fear.txt +4 -0
- package/corpus/prose/wikipedia/First-order_logic.txt +1518 -0
- package/corpus/prose/wikipedia/Fish.txt +10 -0
- package/corpus/prose/wikipedia/Flower.txt +3 -0
- package/corpus/prose/wikipedia/Food.txt +10 -0
- package/corpus/prose/wikipedia/Grass.txt +9 -0
- package/corpus/prose/wikipedia/Hand.txt +2 -0
- package/corpus/prose/wikipedia/Happiness.txt +3 -0
- package/corpus/prose/wikipedia/Heart.txt +4 -0
- package/corpus/prose/wikipedia/Horse.txt +4 -0
- package/corpus/prose/wikipedia/House.txt +6 -0
- package/corpus/prose/wikipedia/Human.txt +4 -0
- package/corpus/prose/wikipedia/Insect.txt +6 -0
- package/corpus/prose/wikipedia/Interactive_fiction.txt +112 -0
- package/corpus/prose/wikipedia/Knowledge.txt +5 -0
- package/corpus/prose/wikipedia/Knowledge_representation_and_reasoning.txt +87 -0
- package/corpus/prose/wikipedia/LICENSE-NOTICE +94 -0
- package/corpus/prose/wikipedia/Language.txt +10 -0
- package/corpus/prose/wikipedia/Learning.txt +4 -0
- package/corpus/prose/wikipedia/Mammal.txt +3 -0
- package/corpus/prose/wikipedia/Memory.txt +5 -0
- package/corpus/prose/wikipedia/Milk.txt +1 -0
- package/corpus/prose/wikipedia/Mountain.txt +1 -0
- package/corpus/prose/wikipedia/Natural_language_processing.txt +211 -0
- package/corpus/prose/wikipedia/Ostrich.txt +2 -0
- package/corpus/prose/wikipedia/Owl.txt +2 -0
- package/corpus/prose/wikipedia/Penguin.txt +2 -0
- package/corpus/prose/wikipedia/Plant.txt +5 -0
- package/corpus/prose/wikipedia/Rain.txt +1 -0
- package/corpus/prose/wikipedia/Resource_Description_Framework.txt +184 -0
- package/corpus/prose/wikipedia/River.txt +1 -0
- package/corpus/prose/wikipedia/School.txt +8 -0
- package/corpus/prose/wikipedia/Sea.txt +1 -0
- package/corpus/prose/wikipedia/Semantic_Web.txt +114 -0
- package/corpus/prose/wikipedia/Semantic_reasoner.txt +29 -0
- package/corpus/prose/wikipedia/Snow.txt +5 -0
- package/corpus/prose/wikipedia/Sun.txt +5 -0
- package/corpus/prose/wikipedia/Teacher.txt +4 -0
- package/corpus/prose/wikipedia/Team.txt +3 -0
- package/corpus/prose/wikipedia/Text-based_game.txt +17 -0
- package/corpus/prose/wikipedia/Tool.txt +4 -0
- package/corpus/prose/wikipedia/Tree.txt +7 -0
- package/corpus/prose/wikipedia/Weather.txt +4 -0
- package/corpus/prose/wikipedia/Web_Ontology_Language.txt +133 -0
- package/corpus/prose/wikipedia/Wind.txt +8 -0
- package/corpus/prose/wikipedia/Writing.txt +5 -0
- package/package.json +2 -1
|
@@ -0,0 +1,211 @@
|
|
|
1
|
+
Natural language processing (NLP) is the processing of natural language information by a computer. NLP is a subfield of computer science and is closely associated with artificial intelligence. NLP is also related to information retrieval, knowledge representation, computational linguistics, and linguistics more broadly.
|
|
2
|
+
Major processing tasks in an NLP system include: speech recognition, text classification, natural language understanding, and natural language generation.
|
|
3
|
+
Natural language processing has its roots in the 1950s. Already in 1950, Alan Turing published an article titled "Computing Machinery and Intelligence," which proposed what is now called the Turing test as a criterion of intelligence, though at the time that was not articulated as a problem separate from artificial intelligence. The proposed test includes a task that involves the automated interpretation and generation of natural language.
|
|
4
|
+
The premise of symbolic NLP is often illustrated using John Searle's Chinese room thought experiment: Given a collection of rules (e.g., a Chinese phrasebook, with questions and matching answers), the computer emulates natural language understanding (or other NLP tasks) by applying those rules to the data it confronts.
|
|
5
|
+
1950s: The Georgetown experiment in 1954 involved fully automatic translation of more than sixty Russian sentences into English. The authors claimed that within three or five years, machine translation would be a solved problem. However, real progress was much slower, and after the ALPAC report in 1966, which found that ten years of research had failed to fulfill the expectations, funding for machine translation was dramatically reduced. Little further research in machine translation was conducted in America (though some research continued elsewhere, such as Japan and Europe) until the late 1980s when the first statistical machine translation systems were developed.
|
|
6
|
+
1960s: Some notably successful natural language processing systems developed in the 1960s were SHRDLU, a natural language system working in restricted "blocks worlds" with restricted vocabularies, and ELIZA, a simulation of Rogerian psychotherapy, written by Joseph Weizenbaum between 1964 and 1966. Despite using minimal information about human thought or emotion, ELIZA was able to produce interactions that appeared human-like. When the "patient" exceeded the very small knowledge base, ELIZA might provide a generic response, for example, responding to "My head hurts" with "Why do you say your head hurts?". Ross Quillian's successful work on natural language was demonstrated with a vocabulary of only twenty words, because that was all that would fit in a computer memory at the time.
|
|
7
|
+
1970s: During the 1970s, many programmers began to write "conceptual ontologies", which structured real-world information into computer-understandable data. Examples are MARGIE (Schank, 1975), SAM (Cullingford, 1978), PAM (Wilensky, 1978), TaleSpin (Meehan, 1976), QUALM (Lehnert, 1977), Politics (Carbonell, 1979), and Plot Units (Lehnert 1981). During this time, the first chatterbots were written (e.g., PARRY).
|
|
8
|
+
1980s: The 1980s and early 1990s mark the heyday of symbolic methods in NLP. Focus areas of the time included research on rule-based parsing (e.g., the development of HPSG as a computational operationalization of generative grammar), morphology (e.g., two-level morphology), semantics (e.g., Lesk algorithm), reference (e.g., within Centering Theory) and other areas of natural language understanding (e.g., in the Rhetorical Structure Theory). Other lines of research were continued, e.g., the development of chatterbots with Racter and Jabberwacky. An important development (that eventually led to the statistical turn in the 1990s) was the rising importance of quantitative evaluation in this period.
|
|
9
|
+
Up until the 1980s, most natural language processing systems were based on complex sets of hand-written rules. Starting in the late 1980s, however, there was a revolution in natural language processing with the introduction of machine learning algorithms for language processing. This shift was influenced by increasing computational power (see Moore's law) and a decline in the dominance of Chomskyan linguistic theories (e.g. transformational grammar), whose theoretical underpinnings discouraged the sort of corpus linguistics that underlies the machine-learning approach to language processing.
|
|
10
|
+
1990s: Many of the notable early successes in statistical methods in NLP occurred in the field of machine translation, due especially to work at IBM Research, such as IBM alignment models. These systems were able to take advantage of existing multilingual textual corpora that had been produced by the Parliament of Canada and the European Union as a result of laws calling for the translation of all governmental proceedings into all official languages of the corresponding systems of government. However, many systems relied on corpora that were specifically developed for the tasks they were designed to perform. This reliance has been a major limitation to their broader effectiveness and continues to affect similar systems. Consequently, significant research has focused on methods for learning effectively from limited amounts of data.
|
|
11
|
+
2000s: With the growth of the web, increasing amounts of raw (unannotated) language data have become available since the mid-1990s. Research has thus increasingly focused on unsupervised and semi-supervised learning algorithms. Such algorithms can learn from data that has not been hand-annotated with the desired answers or using a combination of annotated and non-annotated data. Generally, this task is much more difficult than supervised learning, and typically produces less accurate results for a given amount of input data. However, large quantities of non-annotated data are available (including, among other things, the entire content of the World Wide Web), which can often make up for the worse efficiency if the algorithm used has a low enough time complexity to be practical.
|
|
12
|
+
2003: word n-gram model, at the time the best statistical algorithm, is outperformed by a multi-layer perceptron (with a single hidden layer and context length of several words, trained on up to 14 million words, by Bengio et al.)
|
|
13
|
+
2010: Tomáš Mikolov (then a PhD student at Brno University of Technology) with co-authors applied a simple recurrent neural network with a single hidden layer to language modeling, and in the following years he went on to develop Word2vec. In the 2010s, representation learning and deep neural network-style (featuring many hidden layers) machine learning methods became widespread in natural language processing. This shift gained momentum due to results showing that such techniques can achieve state-of-the-art results in many natural language tasks, e.g., in language modeling and parsing. This is increasingly important in medicine and healthcare, where NLP helps analyze notes and text in electronic health records that would otherwise be inaccessible for study when seeking to improve care or protect patient privacy.
|
|
14
|
+
Symbolic approach, i.e., the hand-coding of a set of rules for manipulating symbols, coupled with a dictionary lookup, was historically the first approach used both by AI in general and by NLP in particular: such as by writing grammars or devising heuristic rules for stemming.
|
|
15
|
+
Machine learning approaches, which include both statistical and neural networks, on the other hand, have many advantages over the symbolic approach:
|
|
16
|
+
both statistical and neural network methods tend to focus more on the most common cases extracted from a corpus of texts, whereas the rule-based approach needs to provide rules for both rare and common cases equally.
|
|
17
|
+
language models, produced by either statistical or neural network methods, are more robust to both unfamiliar (e.g. containing words or structures that have not been seen before) and erroneous input (e.g. with misspelled words or words accidentally omitted) in comparison to the rule-based systems, which are also more costly to produce.
|
|
18
|
+
the larger such a (probabilistic) language model is, the more accurate it becomes, in contrast to rule-based systems that can gain accuracy only by increasing the amount and complexity of the rules leading to intractability problems.
|
|
19
|
+
Rule-based systems are commonly used:
|
|
20
|
+
when the amount of training data is insufficient to successfully apply machine learning methods, e.g., for the machine translation of low-resource languages such as provided by the Apertium system,
|
|
21
|
+
for preprocessing in NLP pipelines, e.g., tokenization, or
|
|
22
|
+
for post-processing and transforming the output of NLP pipelines, e.g., for knowledge extraction from syntactic parses.
|
|
23
|
+
In the late 1980s and mid-1990s, the statistical approach ended a period of AI winter, which was caused by the inefficiencies of the rule-based approaches.
|
|
24
|
+
The earliest decision trees, producing systems of hard if–then rules, were still very similar to the old rule-based approaches.
|
|
25
|
+
Only the introduction of hidden Markov models, applied to part-of-speech tagging, announced the end of the old rule-based approach.
|
|
26
|
+
A major drawback of statistical methods is that they require elaborate feature engineering. Since 2015, neural network–based methods have increasingly replaced traditional statistical approaches, using semantic networks and word embeddings to capture semantic properties of words.
|
|
27
|
+
Intermediate tasks (e.g., part-of-speech tagging and dependency parsing) are not needed anymore.
|
|
28
|
+
Neural machine translation, based on the then-newly invented sequence-to-sequence transformations, made obsolete the intermediate steps, such as word alignment, previously necessary for statistical machine translation.
|
|
29
|
+
The following is a list of some of the most commonly researched tasks in natural language processing. Some of these tasks have direct real-world applications, while others more commonly serve as subtasks that are used to aid in solving larger tasks.
|
|
30
|
+
Though natural language processing tasks are closely intertwined, they can be subdivided into categories for convenience. A coarse division is given below.
|
|
31
|
+
Optical character recognition (OCR)
|
|
32
|
+
Given an image representing printed text, determine the corresponding text.
|
|
33
|
+
Speech recognition
|
|
34
|
+
Given a sound clip of a person or people speaking, determine the textual representation of the speech. This is the opposite of text to speech and is one of the extremely difficult problems colloquially termed "AI-complete" (see above). In natural speech there are hardly any pauses between successive words, and thus speech segmentation is a necessary subtask of speech recognition (see below). In most spoken languages, the sounds representing successive letters blend into each other in a process termed coarticulation, so the conversion of the analog signal to discrete characters can be a very difficult process. Also, given that words in the same language are spoken by people with different accents, the speech recognition software must be able to recognize the wide variety of input as being identical to each other in terms of its textual equivalent.
|
|
35
|
+
Speech segmentation
|
|
36
|
+
Given a sound clip of a person or people speaking, separate it into words. A subtask of speech recognition and typically grouped with it.
|
|
37
|
+
Text-to-speech
|
|
38
|
+
Given a text, transform those units and produce a spoken representation. Text-to-speech can be used to aid the visually impaired.
|
|
39
|
+
Word segmentation (Tokenization)
|
|
40
|
+
Tokenization is a text-processing technique that divides text into individual words or word fragments. This technique results in two key components: a word index and tokenized text. The word index is a list that maps unique words to specific numerical identifiers, and the tokenized text replaces each word with its corresponding numerical token. These numerical tokens are then used in various deep learning methods.
|
|
41
|
+
For a language like English, this is fairly trivial, since words are usually separated by spaces. However, some written languages like Chinese, Japanese and Thai do not mark word boundaries in such a fashion, and in those languages text segmentation is a significant task requiring knowledge of the vocabulary and morphology of words in the language. Sometimes this process is also used in cases like bag of words (BOW) creation in data mining.
|
|
42
|
+
Lemmatization
|
|
43
|
+
The task of removing inflectional endings only and returning the base dictionary form of a word which is also known as a lemma. Lemmatization is another technique for reducing words to their normalized form. But in this case, the transformation actually uses a dictionary to map words to their actual form.
|
|
44
|
+
Morphological segmentation
|
|
45
|
+
Separate words into individual morphemes and identify the class of the morphemes. The difficulty of this task depends greatly on the complexity of the morphology (i.e., the structure of words) of the language being considered. English has fairly simple morphology, especially inflectional morphology, and thus it is often possible to ignore this task entirely and simply model all possible forms of a word (e.g., "open, opens, opened, opening") as separate words. In languages such as Turkish or Meitei, a highly agglutinated Indian language, however, such an approach is not possible, as each dictionary entry has thousands of possible word forms.
|
|
46
|
+
Part-of-speech tagging
|
|
47
|
+
Given a sentence, determine the part of speech (POS) for each word. Many words, especially common ones, can serve as multiple parts of speech. For example, "book" can be a noun ("the book on the table") or verb ("to book a flight"); "set" can be a noun, verb or adjective; and "out" can be any of at least five different parts of speech.
|
|
48
|
+
Stemming
|
|
49
|
+
The process of reducing inflected (or sometimes derived) words to a base form (e.g., "close" will be the root for "closed", "closing", "close", "closer" etc.). Stemming yields similar results as lemmatization, but does so on grounds of rules, not a dictionary.
|
|
50
|
+
Grammar induction
|
|
51
|
+
Generate a formal grammar that describes a language's syntax.
|
|
52
|
+
Sentence breaking (also known as "sentence boundary disambiguation")
|
|
53
|
+
Given a chunk of text, find the sentence boundaries. Sentence boundaries are often marked by periods or other punctuation marks, but these same characters can serve other purposes (e.g., marking abbreviations).
|
|
54
|
+
Parsing
|
|
55
|
+
Determine the parse tree (grammatical analysis) of a given sentence. The grammar for natural languages is ambiguous and typical sentences have multiple possible analyses: perhaps surprisingly, for a typical sentence there may be thousands of potential parses (most of which will seem completely nonsensical to a human). There are two primary types of parsing: dependency parsing and constituency parsing. Dependency parsing focuses on the relationships between words in a sentence (marking things like primary objects and predicates), whereas constituency parsing focuses on building out the parse tree using a probabilistic context-free grammar (PCFG) (see also stochastic grammar).
|
|
56
|
+
Lexical semantics
|
|
57
|
+
What is the computational meaning of individual words in context?
|
|
58
|
+
Distributional semantics
|
|
59
|
+
How can we learn semantic representations from data?
|
|
60
|
+
Named entity recognition (NER)
|
|
61
|
+
Given a stream of text, determine which items in the text map to proper names, such as people or places, and what the type of each such name is (e.g. person, location, organization). Although capitalization can aid in recognizing named entities in languages such as English, this information cannot aid in determining the type of named entity, and in any case, is often inaccurate or insufficient. For example, the first letter of a sentence is also capitalized, and named entities often span several words, only some of which are capitalized. Furthermore, many other languages in non-Western scripts (e.g. Chinese or Arabic) do not have any capitalization at all, and even languages with capitalization may not consistently use it to distinguish names. For example, German capitalizes all nouns, regardless of whether they are names, and French and Spanish do not capitalize names that serve as adjectives. This task is also referred to as token classification.
|
|
62
|
+
Sentiment analysis (see also Multimodal sentiment analysis)
|
|
63
|
+
Sentiment analysis involves identifying and classifying the emotional tone expressed in text. This technique involves analyzing text to determine whether the expressed sentiment is positive, negative, or neutral. Models for sentiment classification typically utilize inputs such as word n-grams, Term Frequency-Inverse Document Frequency (TF-IDF) features, hand-generated features, or employ deep learning models designed to recognize both long-term and short-term dependencies in text sequences. The applications of sentiment analysis are diverse, extending to tasks such as categorizing customer reviews on various online platforms.
|
|
64
|
+
Terminology extraction
|
|
65
|
+
The goal of terminology extraction is to automatically extract relevant terms from a given corpus.
|
|
66
|
+
Word-sense disambiguation (WSD)
|
|
67
|
+
Many words have more than one meaning; we have to select the meaning which makes the most sense in context. For this problem, we are typically given a list of words and associated word senses, e.g. from a dictionary or an online resource such as WordNet.
|
|
68
|
+
Entity linking
|
|
69
|
+
Many words—typically proper names—refer to named entities; here we have to select the entity (a famous individual, a location, a company, etc.) which is referred to in context.
|
|
70
|
+
Relationship extraction
|
|
71
|
+
Given a chunk of text, identify the relationships among named entities (e.g. who is married to whom).
|
|
72
|
+
Semantic parsing
|
|
73
|
+
Given a piece of text (typically a sentence), produce a formal representation of its semantics, either as a graph (e.g., in AMR parsing) or in accordance with a logical formalism (e.g., in DRT parsing). This challenge typically includes aspects of several more elementary NLP tasks from semantics (e.g., semantic role labelling, word-sense disambiguation) and can be extended to include full-fledged discourse analysis (e.g., discourse analysis, coreference; see Natural language understanding below).
|
|
74
|
+
Semantic role labelling (see also implicit semantic role labelling below)
|
|
75
|
+
Given a single sentence, identify and disambiguate semantic predicates (e.g., verbal frames), then identify and classify the frame elements (semantic roles).
|
|
76
|
+
Coreference resolution
|
|
77
|
+
Given a sentence or larger chunk of text, determine which words ("mentions") refer to the same objects ("entities"). Anaphora resolution is a specific example of this task, and is specifically concerned with matching up pronouns with the nouns or names to which they refer. The more general task of coreference resolution also includes identifying so-called "bridging relationships" involving referring expressions. For example, in a sentence such as "He entered John's house through the front door", "the front door" is a referring expression and the bridging relationship to be identified is the fact that the door being referred to is the front door of John's house (rather than of some other structure that might also be referred to).
|
|
78
|
+
Discourse analysis
|
|
79
|
+
This rubric includes several related tasks. One task is discourse parsing, i.e., identifying the discourse structure of a connected text, i.e. the nature of the discourse relationships between sentences (e.g. elaboration, explanation, contrast). Another possible task is recognizing and classifying the speech acts in a chunk of text (e.g. yes–no question, content question, statement, assertion, etc.).
|
|
80
|
+
Implicit semantic role labelling
|
|
81
|
+
Given a single sentence, identify and disambiguate semantic predicates (e.g., verbal frames) and their explicit semantic roles in the current sentence (see Semantic role labelling above). Then, identify semantic roles that are not explicitly realized in the current sentence, classify them into arguments that are explicitly realized elsewhere in the text and those that are not specified, and resolve the former against the local text. A closely related task is zero anaphora resolution, i.e., the extension of coreference resolution to pro-drop languages.
|
|
82
|
+
Recognizing textual entailment
|
|
83
|
+
Given two text fragments, determine if one being true entails the other, entails the other's negation, or allows the other to be either true or false.
|
|
84
|
+
Topic segmentation and recognition
|
|
85
|
+
Given a chunk of text, separate it into segments each of which is devoted to a topic, and identify the topic of the segment.
|
|
86
|
+
Argument mining
|
|
87
|
+
The goal of argument mining is the automatic extraction and identification of argumentative structures from natural language text with the aid of computer programs. Such argumentative structures include the premise, conclusions, the argument scheme and the relationship between the main and subsidiary argument, or the main and counter-argument within discourse.
|
|
88
|
+
Automatic summarization (text summarization)
|
|
89
|
+
Produce a readable summary of a chunk of text. Often used to provide summaries of the text of a known type, such as research papers, articles in the financial section of a newspaper.
|
|
90
|
+
Grammatical error correction
|
|
91
|
+
Grammatical error detection and correction involves a great bandwidth of problems on all levels of linguistic analysis (phonology/orthography, morphology, syntax, semantics, pragmatics). Grammatical error correction is impactful since it affects hundreds of millions of people who use or acquire English as a second language. It has thus been subject to a number of shared tasks since 2011. As far as orthography, morphology, syntax and certain aspects of semantics are concerned, and due to the development of powerful neural language models such as GPT-2, this can now (2019) be considered a largely solved problem and is being marketed in various commercial applications.
|
|
92
|
+
Logic translation
|
|
93
|
+
Translate a text from a natural language into formal logic.
|
|
94
|
+
Machine translation (MT)
|
|
95
|
+
Automatically translate text from one human language to another. This is one of the most difficult problems, and is a member of a class of problems colloquially termed "AI-complete", i.e. requiring all of the different types of knowledge that humans possess (grammar, semantics, facts about the real world, etc.) to solve properly.
|
|
96
|
+
Natural language understanding (NLU)
|
|
97
|
+
Convert chunks of text into more formal representations such as first-order logic structures that are easier for computer programs to manipulate. Natural language understanding involves the identification of the intended semantic from the multiple possible semantics that can be derived from a natural language expression which usually takes the form of organized notations of natural language concepts. Introduction and creation of language metamodels and ontologies are efficient but empirical solutions. An explicit formalization of natural language semantics without confusion with implicit assumptions such as closed-world assumption (CWA) vs. open-world assumption, or subjective Yes/No vs. objective True/False is expected for the construction of a basis of semantics formalization.
|
|
98
|
+
Natural language generation (NLG):
|
|
99
|
+
Convert information from computer databases or semantic intents into readable human language.
|
|
100
|
+
Book generation
|
|
101
|
+
Not an NLP task proper but an extension of natural language generation and other NLP tasks is the creation of full-fledged books. The first machine-generated book was created by a rule-based system in 1984 (Racter, The policeman's beard is half-constructed). The first published work by a neural network was published in 2018, 1 the Road, marketed as a novel, contains sixty million words. Both these systems are basically elaborate but nonsensical (semantics-free) language models. The first machine-generated science book was published in 2019 (Beta Writer, Lithium-Ion Batteries, Springer, Cham). Unlike Racter and 1 the Road, this is grounded on factual knowledge and based on text summarization.
|
|
102
|
+
Document AI
|
|
103
|
+
A Document AI platform sits on top of the NLP technology enabling users with no prior experience of artificial intelligence, machine learning or NLP to quickly train a computer to extract the specific data they need from different document types. NLP-powered Document AI enables non-technical teams to quickly access information hidden in documents, for example, lawyers, business analysts and accountants.
|
|
104
|
+
Dialogue management
|
|
105
|
+
Computer systems intended to converse with a human.
|
|
106
|
+
Question answering
|
|
107
|
+
Given a human-language question, determine its answer. Typical questions have a specific right answer (such as "What is the capital of Canada?"), but sometimes open-ended questions are also considered (such as "What is the meaning of life?").
|
|
108
|
+
Text-to-image generation
|
|
109
|
+
Given a description of an image, generate an image that matches the description.
|
|
110
|
+
Text-to-scene generation
|
|
111
|
+
Given a description of a scene, generate a 3D model of the scene.
|
|
112
|
+
Text-to-video
|
|
113
|
+
Given a description of a video, generate a video that matches the description.
|
|
114
|
+
Based on long-standing trends in the field, it is possible to extrapolate future directions of NLP. As of 2020, three trends among the topics of the long-standing series of CoNLL Shared Tasks can be observed:
|
|
115
|
+
Interest in increasingly abstract, "cognitive" aspects of natural language (1999–2001: shallow parsing, 2002–03: named entity recognition, 2006–09/2017–18: dependency syntax, 2004–05/2008–09 semantic role labelling, 2011–12 coreference, 2015–16: discourse parsing, 2019: semantic parsing).
|
|
116
|
+
Increasing interest in multilinguality, and, potentially, multimodality (English since 1999; Spanish, Dutch since 2002; German since 2003; Bulgarian, Danish, Japanese, Portuguese, Slovenian, Swedish, Turkish since 2006; Basque, Catalan, Chinese, Greek, Hungarian, Italian, Turkish since 2007; Czech since 2009; Arabic since 2012; 2017: 40+ languages; 2018: 60+/100+ languages)
|
|
117
|
+
Elimination of symbolic representations (rule-based over supervised towards weakly supervised methods, representation learning and end-to-end systems)
|
|
118
|
+
Most higher-level NLP applications involve aspects that emulate intelligent behavior and apparent comprehension of natural language. More broadly speaking, the technical operationalization of increasingly advanced aspects of cognitive behavior represents one of the developmental trajectories of NLP(see trends among CoNLL shared tasks above).
|
|
119
|
+
Cognition refers to "the mental action or process of acquiring knowledge and understanding through thought, experience, and the senses." Cognitive science is the interdisciplinary, scientific study of the mind and its processes. Cognitive linguistics is an interdisciplinary branch of linguistics, combining knowledge and research from both psychology and linguistics. Especially during the age of symbolic NLP, the area of computational linguistics maintained strong ties with cognitive studies.
|
|
120
|
+
As an example, George Lakoff offers a methodology to build natural language processing (NLP) algorithms through the perspective of cognitive science, along with the findings of cognitive linguistics, with two defining aspects:
|
|
121
|
+
Apply the theory of conceptual metaphor, explained by Lakoff as "the understanding of one idea, in terms of another" which provides an idea of the intent of the author. For example, consider the English word big. When used in a comparison ("That is a big tree"), the author intends to imply that the tree is physically large relative to other trees or the author's experience. When used metaphorically ("Tomorrow is a big day"), the author intends to imply importance. The intent behind other usages, like in "She is a big person", will remain somewhat ambiguous to a person and a cognitive NLP algorithm alike without additional information.
|
|
122
|
+
Assign relative measures of meaning to a word, phrase, sentence or piece of text based on the information presented before and after the piece of text being analyzed, e.g., by means of a probabilistic context-free grammar (PCFG). The mathematical equation for such algorithms is presented in US Patent 9269353:
|
|
123
|
+
R
|
|
124
|
+
M
|
|
125
|
+
M
|
|
126
|
+
(
|
|
127
|
+
t
|
|
128
|
+
o
|
|
129
|
+
k
|
|
130
|
+
e
|
|
131
|
+
n
|
|
132
|
+
N
|
|
133
|
+
)
|
|
134
|
+
=
|
|
135
|
+
P
|
|
136
|
+
M
|
|
137
|
+
M
|
|
138
|
+
(
|
|
139
|
+
t
|
|
140
|
+
o
|
|
141
|
+
k
|
|
142
|
+
e
|
|
143
|
+
n
|
|
144
|
+
N
|
|
145
|
+
)
|
|
146
|
+
×
|
|
147
|
+
1
|
|
148
|
+
2
|
|
149
|
+
d
|
|
150
|
+
(
|
|
151
|
+
∑
|
|
152
|
+
i
|
|
153
|
+
=
|
|
154
|
+
−
|
|
155
|
+
d
|
|
156
|
+
d
|
|
157
|
+
(
|
|
158
|
+
(
|
|
159
|
+
P
|
|
160
|
+
M
|
|
161
|
+
M
|
|
162
|
+
(
|
|
163
|
+
t
|
|
164
|
+
o
|
|
165
|
+
k
|
|
166
|
+
e
|
|
167
|
+
n
|
|
168
|
+
N
|
|
169
|
+
)
|
|
170
|
+
×
|
|
171
|
+
P
|
|
172
|
+
F
|
|
173
|
+
(
|
|
174
|
+
t
|
|
175
|
+
o
|
|
176
|
+
k
|
|
177
|
+
e
|
|
178
|
+
n
|
|
179
|
+
N
|
|
180
|
+
−
|
|
181
|
+
i
|
|
182
|
+
,
|
|
183
|
+
t
|
|
184
|
+
o
|
|
185
|
+
k
|
|
186
|
+
e
|
|
187
|
+
n
|
|
188
|
+
N
|
|
189
|
+
,
|
|
190
|
+
t
|
|
191
|
+
o
|
|
192
|
+
k
|
|
193
|
+
e
|
|
194
|
+
n
|
|
195
|
+
N
|
|
196
|
+
+
|
|
197
|
+
i
|
|
198
|
+
)
|
|
199
|
+
)
|
|
200
|
+
i
|
|
201
|
+
)
|
|
202
|
+
{\displaystyle {RMM(token_{N})}={PMM(token_{N})}\times {\frac {1}{2d}}\left(\sum _{i=-d}^{d}{((PMM(token_{N})}\times {PF(token_{N-i},token_{N},token_{N+i}))_{i}}\right)}
|
|
203
|
+
Where
|
|
204
|
+
RMM is the relative measure of meaning
|
|
205
|
+
token is any block of text, sentence, phrase or word
|
|
206
|
+
N is the number of tokens being analyzed
|
|
207
|
+
PMM is the probable measure of meaning based on a corpora
|
|
208
|
+
d is the non-zero location of the token along the sequence of N tokens
|
|
209
|
+
PF is the probability function specific to a language
|
|
210
|
+
Ties with cognitive linguistics are part of the historical heritage of NLP, but they have been less frequently addressed since the statistical turn during the 1990s. Nevertheless, approaches to develop cognitive models towards technically operationalizable frameworks have been pursued in the context of various frameworks, e.g., of cognitive grammar, functional grammar, construction grammar, computational psycholinguistics and cognitive neuroscience (e.g., ACT-R), however, with limited uptake in mainstream NLP (as measured by presence on major conferences of the ACL). More recently, ideas of cognitive NLP have been revived as an approach to achieve explainability, e.g., under the notion of "cognitive AI". Likewise, ideas of cognitive NLP are inherent to neural models multimodal NLP (although rarely made explicit) and developments in artificial intelligence, specifically tools and technologies using large language model approaches and new directions in artificial general intelligence based on the free energy principle by British neuroscientist and theoretician at University College London Karl J. Friston.
|
|
211
|
+
Media related to Natural language processing at Wikimedia Commons
|
|
@@ -0,0 +1,2 @@
|
|
|
1
|
+
The ostrich (Struthio camelus) is a large flightless bird that lives in Africa. They are the largest living bird species, and have the biggest eggs of all living birds. Ostriches do not fly, but can run faster than any other bird.
|
|
2
|
+
They are ratites, a useful grouping of medium to large flightless birds. Ostriches have the biggest eyes of all land animals.
|
|
@@ -0,0 +1,2 @@
|
|
|
1
|
+
Owls are birds in the order Strigiformes. There are 200 species, and they are all animals of prey. Most of them are solitary and nocturnal; in fact, they are the only large group of birds which hunt at night. Owls are specialists night-time hunters. They feed on small mammals such as rodents, insects, and other birds, and a few species like to eat fish as well.
|
|
2
|
+
As a group, owls are very successful. They are found in all parts of the world except Antarctica, most of Greenland, and some small islands.
|
|
@@ -0,0 +1,2 @@
|
|
|
1
|
+
Penguins are seabirds in the family Spheniscidae. They use their flippers to swim underwater, but they cannot fly in the air. They eat fish and other seafood. Penguins lay their eggs and raise their babies on land.
|
|
2
|
+
Penguins in the wild only live in the Southern Hemisphere: Antarctica, New Zealand, Australia, South Africa and South America, except Galápagos penguins. The furthest north they get is the Galapagos Islands, where the cold Humboldt Current flows past.
|
|
@@ -0,0 +1,5 @@
|
|
|
1
|
+
Plants are one of six big groups (kingdoms) of living things. They are autotrophic eukaryotes, this means they have complex cells, and make their own food. Usually, they cannot move (not counting growth). Plants need sunlight, soil and water whereas seeds need warmth.
|
|
2
|
+
Plants include familiar types such as trees, herbs, bushes, grasses, vines, ferns, mosses, and green algae. The scientific study of plants, known as botany, has identified about 391,000 extant (living) species of plants.
|
|
3
|
+
Most plants grow in the ground, with stems in the air and roots below the surface. Some float on water. The root part absorbs water and some nutrients the plant needs to live and grow. These climb the stem and reach the leaves. The evaporation of water from pores in the leaves pulls water through the plant. This is called transpiration.
|
|
4
|
+
A plant needs sunlight, carbon dioxide, minerals from the soil and water to make food by photosynthesis. A green substance in plants called chlorophyll traps the energy from the Sun needed to make food. Chlorophyll is mostly found in leaves, inside plastids, which are inside the leaf cells. The leaf can be thought of as a food factory. Leaves of plants vary in shape and size, but they are always the plant organ best suited to capture solar energy. Once the food is made in the leaf, it is transported to the other parts of the plant such as stems and roots.
|
|
5
|
+
The word "plant" can also mean the action of putting something in the ground. For example, farmers plant seeds in the field.
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
Rain is a kind of precipitation. Precipitation is any kind of water that falls from clouds in the sky, like rain, hail, sleet and snow. It is measured by a rain gauge.
|
|
@@ -0,0 +1,184 @@
|
|
|
1
|
+
The Resource Description Framework (RDF) is a method to describe and exchange graph data. It was originally designed as a data model for metadata by the World Wide Web Consortium (W3C). It provides a variety of syntax notations and formats, of which the most widely used is Turtle (Terse RDF Triple Language).
|
|
2
|
+
RDF is a directed graph composed of triple statements. An RDF graph statement is represented by: (1) a node for the subject, (2) an arc from subject to object, representing a predicate, and (3) a node for the object. Each of these parts can be identified by a Internationalized Resource Identifier (IRI). An object can also be a literal value. This simple, flexible data model has a lot of expressive power to represent complex situations, relationships, and other things of interest, while also being appropriately abstract.
|
|
3
|
+
RDF was adopted as a W3C recommendation in 1999. The RDF 1.0 specification was published in 2004, and the RDF 1.1 specification in 2014. SPARQL is a standard query language for RDF graphs. RDF Schema (RDFS), Web Ontology Language (OWL) and SHACL (Shapes Constraint Language) are ontology languages that are used to describe RDF data.
|
|
4
|
+
The RDF data model is similar to classical conceptual modeling approaches (such as entity–relationship or class diagrams). It is based on the idea of making statements about resources (in particular web resources) in expressions of the form subject–predicate–object, known as triples. The subject denotes the resource; the predicate denotes traits or aspects of the resource, and expresses a relationship between the subject and the object.
|
|
5
|
+
For example, one way to represent the notion "The sky has the color blue" in RDF is as the triple: a subject denoting "the sky", a predicate denoting "has the color", and an object denoting "blue". Therefore, RDF uses subject instead of object (or entity) in contrast to the typical approach of an entity–attribute–value model in object-oriented design: entity (sky), attribute (color), and value (blue).
|
|
6
|
+
RDF is an abstract model with several serialization formats (being essentially specialized file formats). In addition the particular encoding for resources or triples can vary from format to format.
|
|
7
|
+
This mechanism for describing resources is a major component in the W3C's Semantic Web activity: an evolutionary stage of the World Wide Web in which automated software can store, exchange, and use machine-readable information distributed throughout the Web, in turn enabling users to deal with the information with greater efficiency and certainty. RDF's simple data model and ability to model disparate, abstract concepts has also led to its increasing use in knowledge management applications unrelated to Semantic Web activity.
|
|
8
|
+
A collection of RDF statements intrinsically represents a labeled, directed multigraph. This makes an RDF data model better suited to certain kinds of knowledge representation than other relational or ontological models.
|
|
9
|
+
As RDFS, OWL and SHACL demonstrate, one can build additional ontology languages upon RDF.
|
|
10
|
+
The initial RDF design, intended to "build a vendor-neutral and operating system- independent system of metadata", derived from the W3C's Platform for Internet Content Selection (PICS), an early web content labelling system, but the project was also shaped by ideas from Dublin Core, and from the Meta Content Framework (MCF), which had been developed during 1995 to 1997 by Ramanathan V. Guha at Apple and Tim Bray at Netscape.
|
|
11
|
+
A first public draft of RDF appeared in October 1997, issued by a W3C working group that included representatives from IBM, Microsoft, Netscape, Nokia, Reuters, SoftQuad, and the University of Michigan.
|
|
12
|
+
In 1999, the W3C published the first recommended RDF specification, the Model and Syntax Specification ("RDF M&S"). This described RDF's data model and an XML serialization.
|
|
13
|
+
Two persistent misunderstandings about RDF developed at this time: firstly, due to the MCF influence and the RDF "Resource Description" initialism, the idea that RDF was specifically for use in representing metadata; secondly that RDF was an XML format rather than a data model, and only the RDF/XML serialisation being XML-based. RDF saw little take-up in this period, but there was significant work done in Bristol, around ILRT at Bristol University and HP Labs, and in Boston at MIT. RSS 1.0 and FOAF became exemplar applications for RDF in this period.
|
|
14
|
+
The recommendation of 1999 was replaced in 2004 by a set of six specifications: "The RDF Primer", "RDF Concepts and Abstract", "RDF/XML Syntax Specification (revised)", "RDF Semantics", "RDF Vocabulary Description Language 1.0", and "The RDF Test Cases".
|
|
15
|
+
This series was superseded in 2014 by the following six "RDF 1.1" documents: "RDF 1.1 Primer", "RDF 1.1 Concepts and Abstract Syntax", "RDF 1.1 XML Syntax", "RDF 1.1 Semantics", "RDF Schema 1.1", and "RDF 1.1 Test Cases".
|
|
16
|
+
The vocabulary defined by the RDF specification is as follows:
|
|
17
|
+
rdf:XMLLiteral
|
|
18
|
+
the class of XML literal values
|
|
19
|
+
rdf:Property
|
|
20
|
+
the class of properties
|
|
21
|
+
rdf:Statement
|
|
22
|
+
the class of RDF statements
|
|
23
|
+
rdf:Alt, rdf:Bag, rdf:Seq
|
|
24
|
+
containers of alternatives, unordered containers, and ordered containers (rdfs:Container is a super-class of the three)
|
|
25
|
+
rdf:List
|
|
26
|
+
the class of RDF Lists
|
|
27
|
+
rdf:nil
|
|
28
|
+
an instance of rdf:List representing the empty list
|
|
29
|
+
rdfs:Resource
|
|
30
|
+
the class resource, everything
|
|
31
|
+
rdfs:Literal
|
|
32
|
+
the class of literal values, e.g. strings and integers
|
|
33
|
+
rdfs:Class
|
|
34
|
+
the class of classes
|
|
35
|
+
rdfs:Datatype
|
|
36
|
+
the class of RDF datatypes
|
|
37
|
+
rdfs:Container
|
|
38
|
+
the class of RDF containers
|
|
39
|
+
rdfs:ContainerMembershipProperty
|
|
40
|
+
the class of container membership properties, rdf:_1, rdf:_2, ..., all of which are sub-properties of rdfs:member
|
|
41
|
+
rdf:type
|
|
42
|
+
an instance of rdf:Property used to state that a resource is an instance of a class
|
|
43
|
+
rdf:first
|
|
44
|
+
the first item in the subject RDF list
|
|
45
|
+
rdf:rest
|
|
46
|
+
the rest of the subject RDF list after rdf:first
|
|
47
|
+
rdf:value
|
|
48
|
+
idiomatic property used for structured values
|
|
49
|
+
rdf:subject
|
|
50
|
+
the subject of the RDF statement
|
|
51
|
+
rdf:predicate
|
|
52
|
+
the predicate of the RDF statement
|
|
53
|
+
rdf:object
|
|
54
|
+
the object of the RDF statement
|
|
55
|
+
rdf:Statement, rdf:subject, rdf:predicate, rdf:object are used for reification (see below).
|
|
56
|
+
rdfs:subClassOf
|
|
57
|
+
the subject is a subclass of a class
|
|
58
|
+
rdfs:subPropertyOf
|
|
59
|
+
the subject is a subproperty of a property
|
|
60
|
+
rdfs:domain
|
|
61
|
+
a domain of the subject property
|
|
62
|
+
rdfs:range
|
|
63
|
+
a range of the subject property
|
|
64
|
+
rdfs:label
|
|
65
|
+
a human-readable name for the subject
|
|
66
|
+
rdfs:comment
|
|
67
|
+
a description of the subject resource
|
|
68
|
+
rdfs:member
|
|
69
|
+
a member of the subject resource
|
|
70
|
+
rdfs:seeAlso
|
|
71
|
+
further information about the subject resource
|
|
72
|
+
rdfs:isDefinedBy
|
|
73
|
+
the definition of the subject resource
|
|
74
|
+
This vocabulary is used as a foundation for RDF Schema, where it is extended.
|
|
75
|
+
Several common serialization formats are in use, including:
|
|
76
|
+
Turtle, a compact, human-friendly format.
|
|
77
|
+
TriG, an extension of Turtle to datasets.
|
|
78
|
+
N-Triples, a very simple, easy-to-parse, line-based format that is not as compact as Turtle.
|
|
79
|
+
N-Quads, a superset of N-Triples, for serializing multiple RDF graphs.
|
|
80
|
+
JSON-LD, a JSON-based serialization.
|
|
81
|
+
N3 or Notation3, a non-standard serialization that is very similar to Turtle, but has some additional features, such as the ability to define inference rules.
|
|
82
|
+
RDF/XML, an XML-based syntax that was the first standard format for serializing RDF.
|
|
83
|
+
RDF/JSON, an alternative syntax for expressing RDF triples using a simple JSON notation.
|
|
84
|
+
RDF/XML is sometimes misleadingly called simply RDF because it was introduced among the other W3C specifications defining RDF and it was historically the first W3C standard RDF serialization format. However, it is important to distinguish the RDF/XML format from the abstract RDF model itself. Although the RDF/XML format is still in use, other RDF serializations are now preferred by many RDF users, both because they are more human-friendly, and because some RDF graphs are not representable in RDF/XML due to restrictions on the syntax of XML QNames.
|
|
85
|
+
With a little effort, virtually any arbitrary XML may also be interpreted as RDF using GRDDL (pronounced 'griddle'), Gleaning Resource Descriptions from Dialects of Languages.
|
|
86
|
+
RDF triples may be stored in a type of database called a triplestore.
|
|
87
|
+
The subject of an RDF statement is either a uniform resource identifier (URI) or a blank node, both of which denote resources. Resources indicated by blank nodes are called anonymous resources. They are not directly identifiable from the RDF statement. The predicate is a URI which also indicates a resource, representing a relationship. The object is a URI, blank node or a Unicode string literal.
|
|
88
|
+
As of RDF 1.1 resources are identified by Internationalized Resource Identifiers (IRIs); IRIs are a generalization of URIs.
|
|
89
|
+
In Semantic Web applications, and in relatively popular applications of RDF like RSS and FOAF (Friend of a Friend), resources tend to be represented by URIs that intentionally denote, and can be used to access, actual data on the World Wide Web. But RDF, in general, is not limited to the description of Internet-based resources. In fact, the URI that names a resource does not have to be dereferenceable at all. For example, a URI that begins with "http:" and is used as the subject of an RDF statement does not necessarily have to represent a resource that is accessible via HTTP, nor does it need to represent a tangible, network-accessible resource—such a URI could represent absolutely anything. However, there is broad agreement that a bare URI (without a # symbol) which returns a 300-level coded response when used in an HTTP GET request should be treated as denoting the internet resource that it succeeds in accessing.
|
|
90
|
+
Therefore, producers and consumers of RDF statements must agree on the semantics of resource identifiers. Such agreement is not inherent to RDF itself, although there are some controlled vocabularies in common use, such as Dublin Core Metadata, which is partially mapped to a URI space for use in RDF. The intent of publishing RDF-based ontologies on the Web is often to establish, or circumscribe, the intended meanings of the resource identifiers used to express data in RDF. For example, the URI:
|
|
91
|
+
http://www.w3.org/TR/2004/REC-owl-guide-20040210/wine#Merlot
|
|
92
|
+
is intended by its owners to refer to the class of all Merlot red wines by vintner (i.e., instances of the above URI each represent the class of all wine produced by a single vintner), a definition which is expressed by the OWL ontology—itself an RDF document—in which it occurs. Without careful analysis of the definition, one might erroneously conclude that an instance of the above URI was something physical, instead of a type of wine.
|
|
93
|
+
Note that this is not a 'bare' resource identifier, but is rather a URI reference, containing the '#' character and ending with a fragment identifier.
|
|
94
|
+
The body of knowledge modeled by a collection of statements may be subjected to reification, in which each statement (that is each triple subject-predicate-object altogether) is assigned a URI and treated as a resource about which additional statements can be made, as in "Jane says that John is the author of document X". Reification is sometimes important in order to deduce a level of confidence or degree of usefulness for each statement.
|
|
95
|
+
In a reified RDF database, each original statement, being a resource, itself, most likely has at least three additional statements made about it: one to assert that its subject is some resource, one to assert that its predicate is some resource, and one to assert that its object is some resource or literal. More statements about the original statement may also exist, depending on the application's needs.
|
|
96
|
+
Borrowing from concepts available in logic (and as illustrated in graphical notations such as conceptual graphs and topic maps), some RDF model implementations acknowledge that it is sometimes useful to group statements according to different criteria, called situations, contexts, or scopes, as discussed in articles by RDF specification co-editor Graham Klyne. For example, a statement can be associated with a context, named by a URI, in order to assert an "is true in" relationship. As another example, it is sometimes convenient to group statements by their source, which can be identified by a URI, such as the URI of a particular RDF/XML document. Then, when updates are made to the source, corresponding statements can be changed in the model, as well.
|
|
97
|
+
Implementation of scopes does not necessarily require fully reified statements. Some implementations allow a single scope identifier to be associated with a statement that has not been assigned a URI, itself. Likewise named graphs in which a set of triples is named by a URI can represent context without the need to reify the triples.
|
|
98
|
+
The predominant query language for RDF graphs is SPARQL. SPARQL is an SQL-like language, and a recommendation of the W3C as of January 15, 2008.
|
|
99
|
+
The following is an example of a SPARQL query to show country capitals in Africa, using a fictional ontology:
|
|
100
|
+
Other non-standard ways to query RDF graphs include:
|
|
101
|
+
RDQL, precursor to SPARQL, SQL-like
|
|
102
|
+
Versa, compact syntax (non–SQL-like), solely implemented in 4Suite (Python).
|
|
103
|
+
RQL, one of the first declarative languages for uniformly querying RDF schemas and resource descriptions, implemented in RDFSuite.
|
|
104
|
+
SeRQL, part of Sesame
|
|
105
|
+
XUL has a template element in which to declare rules for matching data in RDF. XUL uses RDF extensively for data binding.
|
|
106
|
+
SHACL Advanced Features specification (W3C Working Group Note), the most recent version of which is maintained by the SHACL Community Group, defines support for SHACL Rules, used for data transformations, inferences and mappings of RDF based on SHACL shapes.
|
|
107
|
+
The predominant language for describing and validating RDF graphs is SHACL (Shapes Constraint Language). SHACL specification is divided in two parts: SHACL Core and SHACL-SPARQL. SHACL Core consists of a list of built-in constraints such as cardinality, range of values and many others. SHACL-SPARQL describes SPARQL-based constraints and an extension mechanism to declare new constraint components.
|
|
108
|
+
Other ways to describe and validate RDF graphs include:
|
|
109
|
+
SPARQL Inferencing Notation (SPIN) is a de-facto standard vocabulary for metamodeling SPARQL queries in RDF. SPIN provides semantics for defining inferencing rules and constraints for validation; however, these features have been superseded by SHACL.
|
|
110
|
+
ShEx (Shape Expressions) is a concise language for RDF validation and description.
|
|
111
|
+
The following example is taken from the W3C website describing a resource with statements "there is a Person identified by http://www.w3.org/People/EM/contact#me, whose name is Eric Miller, whose email address is e.miller123(at)example (changed for security purposes), and whose title is Dr."
|
|
112
|
+
The resource "http://www.w3.org/People/EM/contact#me" is the subject.
|
|
113
|
+
The objects are:
|
|
114
|
+
"Eric Miller" (with a predicate "whose name is"),
|
|
115
|
+
mailto:e.miller123(at)example (with a predicate "whose email address is"), and
|
|
116
|
+
"Dr." (with a predicate "whose title is").
|
|
117
|
+
The subject is a URI.
|
|
118
|
+
The predicates also have URIs. For example, the URI for each predicate:
|
|
119
|
+
"whose name is" is http://www.w3.org/2000/10/swap/pim/contact#fullName,
|
|
120
|
+
"whose email address is" is http://www.w3.org/2000/10/swap/pim/contact#mailbox,
|
|
121
|
+
"whose title is" is http://www.w3.org/2000/10/swap/pim/contact#personalTitle.
|
|
122
|
+
In addition, the subject has a type (with URI http://www.w3.org/1999/02/22-rdf-syntax-ns#type), which is person (with URI http://www.w3.org/2000/10/swap/pim/contact#Person).
|
|
123
|
+
Therefore, the following "subject, predicate, object" RDF triples can be expressed:
|
|
124
|
+
http://www.w3.org/People/EM/contact#me, http://www.w3.org/2000/10/swap/pim/contact#fullName, "Eric Miller"
|
|
125
|
+
http://www.w3.org/People/EM/contact#me, http://www.w3.org/2000/10/swap/pim/contact#mailbox, mailto:e.miller123(at)example
|
|
126
|
+
http://www.w3.org/People/EM/contact#me, http://www.w3.org/2000/10/swap/pim/contact#personalTitle, "Dr."
|
|
127
|
+
http://www.w3.org/People/EM/contact#me, http://www.w3.org/1999/02/22-rdf-syntax-ns#type, http://www.w3.org/2000/10/swap/pim/contact#Person
|
|
128
|
+
In standard N-Triples format, this RDF can be written as:
|
|
129
|
+
Equivalently, it can be written in standard Turtle (syntax) format as:
|
|
130
|
+
Or more concisely, using a common shorthand syntax of Turtle as:
|
|
131
|
+
Or, it can be written in RDF/XML format as:
|
|
132
|
+
Certain concepts in RDF are taken from logic and linguistics, where subject-predicate and subject-predicate-object structures have meanings similar to, yet distinct from, the uses of those terms in RDF. This example demonstrates:
|
|
133
|
+
In the English language statement 'New York has the postal abbreviation NY' , 'New York' would be the subject, 'has the postal abbreviation' the predicate and 'NY' the object.
|
|
134
|
+
Encoded as an RDF triple, the subject and predicate would have to be resources named by URIs. The object could be a resource or literal element. For example, in the N-Triples form of RDF, the statement might look like:
|
|
135
|
+
In this example, "urn:x-states:New%20York" is the URI for a resource that denotes the US state New York, "http://purl.org/dc/terms/alternative" is the URI for a predicate (whose human-readable definition can be found here), and "NY" is a literal string. Note that the URIs chosen here are not standard, and do not need to be, as long as their meaning is known to whatever is reading them.
|
|
136
|
+
In a like manner, given that "https://en.wikipedia.org/wiki/Tony_Benn" identifies a particular resource (regardless of whether that URI could be traversed as a hyperlink, or whether the resource is actually the Wikipedia article about Tony Benn), to say that the title of this resource is "Tony Benn" and its publisher is "Wikipedia" would be two assertions that could be expressed as valid RDF statements. In the N-Triples form of RDF, these statements might look like the following:
|
|
137
|
+
To an English-speaking person, the same information could be represented simply as:
|
|
138
|
+
The title of this resource, which is published by Wikipedia, is 'Tony Benn'
|
|
139
|
+
However, RDF puts the information in a formal way that a machine can understand. The purpose of RDF is to provide an encoding and interpretation mechanism so that resources can be described in a way that particular software can understand it; in other words, so that software can access and use information that it otherwise could not use.
|
|
140
|
+
Both versions of the statements above are wordy because one requirement for an RDF resource (as a subject or a predicate) is that it be unique. The subject resource must be unique in an attempt to pinpoint the exact resource being described. The predicate needs to be unique in order to reduce the chance that the idea of Title or Publisher will be ambiguous to software working with the description. If the software recognizes http://purl.org/dc/elements/1.1/title (a specific definition for the concept of a title established by the Dublin Core Metadata Initiative), it will also know that this title is different from a land title or an honorary title or just the letters t-i-t-l-e put together.
|
|
141
|
+
The following example, written in Turtle, shows how such simple claims can be elaborated on, by combining multiple RDF vocabularies. Here, we note that the primary topic of the Wikipedia page is a "Person" whose name is "Tony Benn":
|
|
142
|
+
DBpedia – Extracts facts from Wikipedia articles and publishes them as RDF data.
|
|
143
|
+
YAGO – Similar to DBpedia extracts facts from Wikipedia articles and publishes them as RDF data.
|
|
144
|
+
Wikidata – Collaboratively edited knowledge base hosted by the Wikimedia Foundation.
|
|
145
|
+
Creative Commons – Uses RDF to embed license information in web pages and mp3 files.
|
|
146
|
+
FOAF (Friend of a Friend) – designed to describe people, their interests and interconnections.
|
|
147
|
+
Haystack client – Semantic web browser from MIT CS & AI lab.
|
|
148
|
+
IDEAS Group – developing a formal 4D ontology for Enterprise Architecture using RDF as the encoding.
|
|
149
|
+
Microsoft shipped a product, Connected Services Framework, which provides RDF-based Profile Management capabilities.
|
|
150
|
+
MusicBrainz – Publishes information about Music Albums.
|
|
151
|
+
NEPOMUK, an open-source software specification for a Social Semantic desktop uses RDF as a storage format for collected metadata. NEPOMUK is mostly known because of its integration into the KDE SC 4 desktop environment.
|
|
152
|
+
Cochrane is a global publisher of clinical study meta-analyses in evidence based healthcare. They use an ontology driven data architecture to semantically annotate their published reviews with RDF based structured data.
|
|
153
|
+
RDF Site Summary – one of several "RSS" languages for publishing information about updates made to a web page; it is often used for disseminating news article summaries and sharing weblog content.
|
|
154
|
+
Simple Knowledge Organization System (SKOS) – a KR representation intended to support vocabulary/thesaurus applications
|
|
155
|
+
SIOC (Semantically-Interlinked Online Communities) – designed to describe online communities and to create connections between Internet-based discussions from message boards, weblogs and mailing lists.
|
|
156
|
+
Smart-M3 – provides an infrastructure for using RDF and specifically uses the ontology agnostic nature of RDF to enable heterogeneous mashing-up of information
|
|
157
|
+
LV2 - a libre plugin format using Turtle to describe API/ABI capabilities and properties
|
|
158
|
+
Software Package Data Exchange – A standard for specifying bills of material.
|
|
159
|
+
Some uses of RDF include research into social networking. It will also help people in business fields understand better their relationships with members of industries that could be of use for product placement. It will also help scientists understand how people are connected to one another.
|
|
160
|
+
RDF is being used to gain a better understanding of road traffic patterns. This is because the information regarding traffic patterns is on different websites, and RDF is used to integrate information from different sources on the web. Before, the common methodology was using keyword searching, but this method is problematic because it does not consider synonyms. This is why ontologies are useful in this situation. But one of the issues that comes up when trying to efficiently study traffic is that to fully understand traffic, concepts related to people, streets, and roads must be well understood. Since these are human concepts, they require the addition of fuzzy logic. This is because values that are useful when describing roads, like slipperiness, are not precise concepts and cannot be measured. This would imply that the best solution would incorporate both fuzzy logic and ontology.
|
|
161
|
+
Linked data
|
|
162
|
+
JSON-LD
|
|
163
|
+
Notation3
|
|
164
|
+
RDF/XML
|
|
165
|
+
RDFa
|
|
166
|
+
TriG
|
|
167
|
+
TriX
|
|
168
|
+
RDF Mapping Language (RML)
|
|
169
|
+
Entity–attribute–value model
|
|
170
|
+
Graph theory – an RDF model is a labeled, directed multi-graph.
|
|
171
|
+
SciCrunch
|
|
172
|
+
Semantic network
|
|
173
|
+
Tag (metadata)
|
|
174
|
+
Business Intelligence 2.0 (BI 2.0)
|
|
175
|
+
Data portability
|
|
176
|
+
EU Open Data Portal
|
|
177
|
+
Folksonomy
|
|
178
|
+
RDF Schema
|
|
179
|
+
Semantic technology
|
|
180
|
+
Swoogle
|
|
181
|
+
Universal Networking Language (UNL)
|
|
182
|
+
VoID
|
|
183
|
+
W3C's RDF at W3C: specifications, guides, and resources
|
|
184
|
+
RDF Semantics: specification of semantics, and complete systems of inference rules for both RDF and RDFS
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
A river is a stream of water that flows through a channel on the surface of the ground. The passage where the river flows is called the riverbed and the earth on each side is called a riverbank. A river begins on high ground or in hills or mountains and flows down from the high ground to the lower ground, because of gravity. A river begins as a small stream and gets bigger the farther it flows.
|
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
A school is an educational learning institution designed to provide classes and learning environment to teach students, usually with help of a teacher or professor. Some senior secondary schools can also be considered as a college. Almost all countries and states have systems of formal education, which is sometimes compulsory or legally required. There are various types of schools or learning programs. Topics such as reading, writing, and mathematics are central to education for children.
|
|
2
|
+
Most of a student's time is spent in a classroom. This is where 10 to 30 people sit to take part in educational discussion. In the United States, the average number of students per classroom in primary schools is 23.1.
|
|
3
|
+
The term "school" is used for many educational environments – particularly in North America. In North America, a person taking a first degree at a university is often self-described as "going to school". In Europe that would never be the case. They would describe themselves as "going to university". The style of university education can be so different between countries.
|
|
4
|
+
There are different types of schools: elementary schools (primary in the UK), middle schools (secondary in the UK), and so on.
|
|
5
|
+
In many places around the world, children must go to school for a certain number of years. Learning may take place in the classroom, in outside environments, or on visits to other places. Colleges and universities are places to learn for students over 17 or 18 years of age. Vocational schools teach skills people need for jobs.
|
|
6
|
+
Some people attend school longer than others. This is because some jobs require more training than others, like for example becoming a doctor takes about 10-14 years of education. For young children, one teacher may teach all subjects. Teachers for older students are more specialized, and they only teach a few subjects. Common subjects taught include science, arts such as music, humanities, like geography and history, and languages.
|
|
7
|
+
Children with mental health problems, and problems such as autism and other conditions, usually do not go to regular schools. These children are given other ways to get schooling, like special schools. There also are special schools which teach things which regular schools do not.
|
|
8
|
+
Graduate schools are in universities. They are for students who have graduated with a first degree from colleges and universities. The aim is to offer Masters and PhD courses to the best students.
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
A sea is a large body of salt water. It may be an ocean, or may be a large saltwater lake which like the Caspian Sea, lacks a natural outlet.
|