Glossary

Natural Language Processing (NLP)

The field of AI focused on enabling computers to understand, interpret, and generate human language — the foundation for LLMs, semantic search, text extraction, and most AI research tools.


What it means

Natural language processing (NLP) is the subfield of artificial intelligence concerned with the computational processing of human language — reading, understanding, generating, and translating text. It is the foundational discipline underlying large language models, semantic search, information extraction, machine translation, and automatic summarization.

NLP has been transformed by the transformer architecture and the large language model paradigm since 2017–2020. Most NLP tasks that were previously handled by separate specialized models (named entity recognition, sentiment analysis, question answering, summarization) are now handled by fine-tuning or prompting a single large model.

Core NLP tasks researchers encounter in AI tools:

Task What it does Example in research tools
Named entity recognition (NER) Identifies and classifies entities (genes, chemicals, diseases) in text Extracting compound names from chemistry papers
Text classification Assigns labels to text Rayyan’s inclusion/exclusion classification
Semantic search Retrieves semantically related documents Elicit, Semantic Scholar
Summarization Generates condensed summaries of text NotebookLM, Elicit TLDR
Information extraction Extracts structured data from unstructured text Elicit’s custom extraction fields
Question answering Answers questions given a context document NotebookLM, Claude with PDFs

Why it matters for researchers

All the tools on this site are built on NLP. Understanding what NLP can and cannot do reliably helps set appropriate expectations:

  • NLP excels at tasks where the answer is contained explicitly in the source text
  • NLP struggles with tasks requiring external knowledge, numerical reasoning over tables, and drawing valid inferences that require domain expertise
  • Biomedical NLP is a distinct subfield: general NLP models perform worse on scientific text because the vocabulary, sentence structure, and reasoning patterns differ from the web text they were trained on. Domain-specific models (BioBERT, PubMedBERT, BioGPT) are often better for specialized medical/biological text extraction tasks