Lexical analysis
Computing process of parsing a sequence of characters into a sequence of tokens
Lexical tokenization is conversion of a text into (semantically or syntactically) meaningful lexical tokens belonging to categories defined by a "lexer" program. In case of a natural language, those categories include nouns, verbs, adjectives, punctuations etc. In case of a programming language, the categories include identifiers, operators, grouping symbols, data types and language keywords.
Nº Q835922 ★★
Uncommon · Knowledge
Lexical analysis
Computing process of parsing a sequence of characters into a sequence of tokens
Lexical tokenization is conversion of a text into (semantically or syntactically) meaningful lexical tokens belonging to categories defined by a "lexer" program. In case of a natural language, those categories include nouns, verbs, adjectives, punctuations etc. In case of a programming language, the categories include identifiers, operators, grouping symbols, data types and language keywords.
Last price
—
Floor price
—
7-day median
—
30-day sales
0
30-day range
—
In circulation
0
Price history
median
low – high
sales
No sales in this period
Show table
| Date | median | Low | High | sales |
|---|
Sales history
- Last sale
- —
- 30-day average
- —
- 30-day low
- —
- 30-day high
- —
- Sales 7d
- 0
- Sales 30d
- 0
No sales yet.
Anonymous sales: no buyer or seller shown. Figures count player-to-player sales only.
From Wikipedia
Lexical tokenization is conversion of a text into (semantically or syntactically) meaningful lexical tokens belonging to categories defined by a "lexer" program. In case of a natural language, those categories include nouns, verbs, adjectives, punctuations etc. In case of a programming language, the categories include identifiers, operators, grouping symbols, data types and language keywords. Lexical tokenization is related to the type of tokenization used in large language models (LLMs) but with two differences. First, lexical tokenization is usually based on a lexical grammar, whereas LLM tokenizers are usually probability-based. Second, LLM tokenizers perform a second step that converts the tokens into numerical values.
Text: Wikipédia, CC BY-SA 4.0. ·
Related cards
LL parser
Left-to-right, leftmost derivation top-down parser for a subset of context-free languages
Nº Q932615 ★★
Yacc
Parser generator
Nº Q305932 ★
Recursive descent parser
Style of parser written as a recursive structure matching the grammar it parses
Nº Q1323264 ★★
Assembly language
Any low-level programming language in which there is a very strong correspondence between the instructions in the language and the architecture's machine code instructions
Nº Q165436 ★★★★
Comment (computer programming)
Explanatory note in the source code of a computer program, skipped during its compilation or interpretation and without effect on execution
Nº Q1141067 ★
Interpreter (computing)
Program that executes source code without a separate compilation step
Nº Q183065 ★★★