Category: Productivity

  • Token Parser: What It Is and How to Build One

    Token Parser: What It Is and How to Build One

    Header: one name, four different parsers

    A token parser turns an input string into discrete tokens and interprets them: lexical tokens in source code, Base64url segments in a JWT, placeholders like {SEQ:X} in a document template, or design tokens in a build pipeline. Which parser you need depends on the format you’re decoding — and the term still spans four unrelated meanings that no single page explains.

    What Is a Token Parser?

    A token parser is the component that takes a raw string and produces a sequence of typed, interpreted tokens. It answers three questions, in order: where does one meaningful unit end and the next begin, what kind of unit is it, and what does it mean once you know its type? Everything downstream — an authorization header, an abstract syntax tree, a generated document number, a compiled design system — inherits whatever the parser decided.

    That definition also marks the parser’s boundary. A token parser is not a validator, a verifier, or a serializer. Parsing ends where trust decisions begin. Decoding a JWT’s payload and confirming the token was signed by a party holding the private key are different operations, run by different code, with different failure modes. Conflating the two is the most common defect in this space.

    The four variants in this guide share the same skeleton, even though they operate on completely unrelated formats:

    • A scanner or tokenizer that walks the input and segments it
    • A token grammar — a regex, a DFA, or a lookup registry — that decides which segment matches which token type
    • A token record carrying at minimum a type, the raw lexeme, an interpreted value, and a position for error reporting
    • An error path that defines what happens when the input does not match the grammar

    Shared skeleton of the four token parser variants

    That ambiguity is the real problem. Search for “token” today and you’ll find at least five meanings: a lexeme in source code, a Base64url segment inside a JWT, an OAuth access token presented in an Authorization header, a design token in a build pipeline, a template placeholder — and, in developer tooling, a counted unit of LLM consumption. The premise of the ezparser JWT parser, for example, is that “token” means a signed JSON credential. The premise of the plus-uno token-literal check is that “token” means a design-system value. Both are called token parsers. They don’t read the same thing.

    This matters because picking the wrong branch gives you a parser that compiles, runs, passes a smoke test, and reads the wrong object. Point a JWT-shaped parser at a design-token file and it won’t throw — it will simply extract nothing useful, or extract something plausible but wrong.

    Token parser vs. tokenizer vs. lexer: where the lines are

    These terms overlap, but they describe different stages. A tokenizer or lexer produces the token stream from a character string. A token parser consumes that stream and assigns structural meaning — grouping tokens into expressions, blocks, or records.

    In practice, plenty of codebases use the words interchangeably, and the distinction often collapses inside a single module. The useful rule: check what the library actually accepts and returns before assuming which term it means. If it takes a string and returns a list of typed units, the vocabulary is cosmetic. If it takes a string and returns a tree, it’s doing both jobs, and the naming isn’t worth arguing about.

    Quick glossary: token, lexeme, claim, Base64url, design token, placeholder

    • Token — a typed unit produced by parsing; its type, not its text, drives downstream behavior.
    • Lexeme — the raw substring a token was matched from, kept for error messages and round-tripping.
    • Claim — a name/value pair inside a JWT payload, such as iss, sub, or exp.
    • Base64url — a URL-safe Base64 variant that substitutes - for + and _ for /, and typically omits padding.
    • Design token — a named design value (color, dimension, typography, motion) held in a platform artifact and consumed by a build pipeline.
    • Placeholder — a bracketed token inside a template string, such as {BRANCH} or {SEQ:X}, replaced at generation time.

    Which Token Parser Do You Actually Need? A Decision Table

    The input in front of you determines which parser you should reach for. Find your input in the first column, then read across — the rest of this guide follows this table.

    Input you have Token meaning Output you want Parser type First thing to get right
    A three-part dotted string copied out of a browser or log JWT / OAuth security token Readable header and claims Base64url segment decoder + JSON deserializer Decode only — never verify, never trust for authorization
    Source code, a log line, or a configuration file Lexical token A token stream feeding a parser or AST Tokenizer / lexer Maximal munch and rule ordering
    A string containing {PREFIX} or {SEQ:X} Template placeholder token A substituted, unique output string Registry-driven substitution Fail loud on unknown tokens
    JSON, SCSS, or Kotlin artifacts holding colors, spacing, typography Design token Generated platform code Build-pipeline generator with a pinned source Pin the source commit and reproduce byte for byte

    Every row works on a different contract. The JWT row is inspection only — the output goes to a human debugging a login flow. The lexical row is a pipeline stage. The template row is a substitution engine where correctness means uniqueness, not readability. The design-token row is a code generator where correctness means diff size.

    None of the pages that currently rank for this term defines it or separates those four meanings. They’re either tool landing pages describing one branch in isolation, or GitHub issues that assume you already know which branch you’re on. The table above fills that gap.

    One-line rule: the format you are decoding picks the parser

    If you remember nothing else, remember this: the format you’re decoding picks the parser, not the other way around. You don’t choose a token parser and then go looking for something to feed it. Name the token format first, and the format dictates the grammar, the failure behavior, and where the parser has to stop.

    Token Parser for Security Tokens: How to Decode and Inspect a JWT

    The JWT parser is the most-searched branch of this topic, and the one where misuse causes the most damage. The procedure below is for inspection only. It’s meant for a debugger, a local script, or a CLI — never for a server-side authorization decision.

    Step 1 — split on the dots. RFC 7519 defines a compact JWT as three Base64url segments joined by periods: Header, Payload, and Signature. Splitting gives you three strings. Anything other than three is a malformed token.

    Step 2 — Base64url-decode the first two segments. Base64url substitutes - for + and _ for /, and the padding is usually stripped. A standard Base64 decoder will fail on a stripped segment unless you re-pad it to a multiple of four characters. This is the most common reason a naive decoder throws on a valid token.

    Step 3 — read the fields that matter. In the header, alg tells you which signing algorithm was declared, and typ usually reads JWT. In the payload, iss is the issuer, sub is the subject, aud is the intended audience, iat is issued-at, and exp is expiry — all timestamps in Unix seconds. Convert exp to a UTC time and compare it against the current clock; that answers “is this token stale?” without a round-trip to the issuer.

    Step 4 — treat the signature segment as opaque bytes. Without the HMAC secret or the RSA public key, the third segment tells you nothing. It’s a byte string whose only purpose is to be checked by code that holds the key.

    Step 5 — write the decoder. The following examples are explicitly decode-only.

    import base64, json, datetime
    
    def decode_segment(segment: str) -> dict:
        # Base64url: restore padding before decoding
        padded = segment + "=" * (-len(segment) % 4)
        raw = base64.urlsafe_b64decode(padded.encode("ascii"))
        return json.loads(raw.decode("utf-8"))
    
    def inspect(jwt: str) -> None:
        parts = jwt.split(".")
        if len(parts) != 3:
            raise ValueError("compact JWT must have exactly three segments")
        header, payload, _signature = parts  # signature is opaque here
        header = decode_segment(header)
        payload = decode_segment(payload)
    
        print("alg:", header.get("alg"), "typ:", header.get("typ"))
        for claim in ("iss", "sub", "aud", "iat", "exp"):
            print(f"{claim}: {payload.get(claim)}")
    
        if "exp" in payload:
            expiry = datetime.datetime.fromtimestamp(payload["exp"], datetime.timezone.utc)
            print("expires at (UTC):", expiry.isoformat())
            print("expired:", expiry < datetime.datetime.now(datetime.timezone.utc))
    

    The Node equivalent is shorter because the runtime already understands the encoding:

    // decode-only — this does not verify the signature
    function inspect(jwt) {
      const parts = jwt.split('.');
      if (parts.length !== 3) throw new Error('compact JWT must have exactly three segments');
    
      const [header, payload] = parts.map((segment) =>
        JSON.parse(Buffer.from(segment, 'base64url').toString('utf8'))
      );
    
      console.log('alg:', header.alg, 'typ:', header.typ);
      for (const claim of ['iss', 'sub', 'aud', 'iat', 'exp']) {
        console.log(claim, payload[claim]);
      }
      if (payload.exp) {
        const expiry = new Date(payload.exp * 1000);
        console.log('expires at (UTC):', expiry.toISOString());
        console.log('expired:', expiry < new Date());
      }
    }
    

    Step 6 — place the token in its OAuth context. In OAuth 2.0, an access token arrives in an Authorization: Bearer <access_token> header. The token_type value issued alongside the token is what tells a client how to present it — not an assumption baked into the calling code.

    Three-step JWT decode-and-inspect flow

    Splitting and Base64url-decoding the three segments

    The padding problem deserves its own note, because it produces confusing failures. Base64url encoding of a JSON object rarely lands on a multiple of four characters, and most producers strip the = padding rather than emit it. A decoder that expects padded input either throws or silently truncates. Re-padding with (-len(segment) % 4) characters — as in the Python example above — restores the input to a form the standard decoder will accept.

    The second trap is the alphabet. - and _ aren’t valid characters in standard Base64. A decoder that doesn’t switch alphabets will reject a perfectly valid segment. Python’s base64.urlsafe_b64decode and Node’s 'base64url' encoding handle this automatically; a hand-rolled decoder usually won’t.

    Reading exp, iss, aud and sub without a server round-trip

    Four claims carry most of the debugging value. exp tells you whether the token is still current. iss tells you which system minted it — the usual cause of a rejected token is that it came from a staging issuer when production was expected. aud tells you which API it was meant for, and a token minted for one audience will be correctly rejected by another. sub tells you who the token is about.

    All four can be read offline, with no key required. That’s the point: claims are informative, not authoritative. Anything a decoder can read, an attacker can write.

    Where a Token Parser Must Stop: Decoding Is Not Verification

    Decoding reverses Base64url and deserializes JSON. That’s the whole operation. Anyone can craft a token with any claims they like, sign nothing, and produce a string that decodes perfectly. Only signature verification with the HMAC secret or RSA public key establishes that the claims actually came from the issuer. A token parser that returns claims to a caller who then treats them as trusted isn’t misconfigured — it’s been asked to do something it can’t do.

    The boundary generalizes across credential formats. A SAML response decoder shows every field of an assertion without proving it came from the identity provider, and a cookie header parser shows what the browser sent without vouching for the session behind it. In each case, parsing makes claims readable; only verification makes them trustworthy.

    The boundary line between parsing and trust

    OAuth 2.0 makes that boundary explicit. RFC 6749 §5.1 requires both access_token and token_type in a successful token response, and states that token_type is case-insensitive — so Bearer, bearer, and BEARER are all valid. RFC 6749 §7.1 is the rule that gets skipped in practice: a client must not use an access token whose type it doesn’t understand.

    Reject a missing token_type instead of defaulting to Bearer

    The cortexfs issue #282 documents exactly this failure. In parse_oauth_token_response, validation used Option::is_some_and, which only rejects a present non-Bearer value. A response with no token_type field at all passed validation silently. Downstream, model_headers serialized every Credential::OAuth as Authorization: Bearer <access_token> — so an undeclared type got reinterpreted as Bearer by default.

    The shape of the fix generalizes well beyond one repository. Require the field to be present. Accept Bearer case-insensitively. Keep rejecting unsupported types such as mac. Update the success fixtures in the same change to include token_type: "Bearer". As the issue records, the expected production change was one existing predicate tightened in place, netting zero production lines of code — the fix removed an implicit assumption instead of adding an abstraction.

    Here’s the pattern to internalize: a validator that checks “if present, must be valid” is a weaker thing than one that checks “must be present and valid.” The first admits a missing field; the second doesn’t. When the consuming code has a fallback default, the first becomes a silent policy decision made by omission.

    Keep parser errors bounded and generic

    A token parser must never return its internals to a client. Offsets, grammar rule names, expected token types, partial parse states — that’s a map of the input format, and handing that map to an untrusted caller is a gift.

    The synchro issue #53 records this as audit finding C12, independently verified with a verdict of CONFIRMED at severity Low. The server returned token parser details in responses, and no existing test covered the behavior. The proposed fix is to return a bounded generic error instead of parser internals, with an acceptance criterion that a permanent test fails without the fix and passes with it. Detailed diagnostics belong in server-side logs behind a stable correlation ID — the client gets an opaque error and a handle, not a description of the grammar.

    Building a Lexical Token Parser (Tokenizer or Lexer)

    For source code and structured text, the pipeline runs: raw string → scanner → token stream → parser → AST. The token parser is the stage that decides where one unit ends and the next begins, and every stage after it inherits those decisions.

    Lexical pipeline from raw string to AST

    Define the token record first. A workable minimum has four fields: an enum type, the raw lexeme, an optional interpreted value (a parsed number, an unescaped string), and a line/column position for diagnostics. Skipping the position field is the most common shortcut and the one that hurts most later — error messages without positions are close to useless.

    Three implementation rules prevent most tokenizer bugs:

    1. Maximal munch. Always take the longest valid match at the current position, so == is one token rather than two =.
    2. Order rules so keywords precede identifiers. If identifiers are matched first, while becomes an identifier and the keyword rule never fires.
    3. Decide explicitly about whitespace and comments. Either emit them as token kinds or skip them deliberately — but make it a documented decision, not an accident of the first regex that happens to match.

    A working hand-written lexer fits in roughly forty lines and covers the token kinds most languages need:

    import re
    from dataclasses import dataclass
    
    @dataclass
    class Token:
        type: str
        lexeme: str
        value: object = None
        line: int = 1
        col: int = 1
    
    # Order matters: keywords before identifiers, longest operators first.
    RULES = [
        ("NUMBER", r"\d+(\.\d+)?"),
        ("IDENT",  r"[A-Za-z_]\w*"),
        ("STRING", r'"(?:[^"\\]|\\.)*"'),
        ("OP",     r"==|!=|<=|>=|&&|\|\||[-+*/=<>(){},;]"),
    ]
    KEYWORDS = {"if", "else", "while", "return"}
    SKIP = re.compile(r"[ \t\r]+")
    
    def tokenize(src: str) -> list[Token]:
        tokens, i, line, line_start = [], 0, 1, 0
        while i < len(src):
            if src[i] == "\n":
                line, i, line_start = line + 1, i + 1, i
                continue
            m = SKIP.match(src, i)
            if m:
                i = m.end()
                continue
            for kind, pattern in RULES:
                m = re.compile(pattern).match(src, i)
                if not m:
                    continue
                lexeme = m.group(0)
                if kind == "IDENT" and lexeme in KEYWORDS:
                    kind = "KEYWORD"
                value = None
                if kind == "NUMBER":
                    value = float(lexeme) if "." in lexeme else int(lexeme)
                elif kind == "STRING":
                    value = lexeme[1:-1]
                tokens.append(Token(kind, lexeme, value, line, i - line_start + 1))
                i = m.end()
                break
            else:
                raise SyntaxError(f"unexpected character {src[i]!r} "
                                  f"at line {line}, column {i - line_start + 1}")
        tokens.append(Token("EOF", "", None, line, i - line_start + 1))
        return tokens
    
    if __name__ == "__main__":
        for t in tokenize('total = price * 2; if total > 100 { return "big"; }'):
            print(f"{t.type:<8} {t.lexeme!r:<12} {t.value!r:<8} {t.line}:{t.col}")
    

    Printing the token dump isn’t decoration. The dump is what you compare against when a grammar change produces unexpected behavior, and it separates “the lexer is wrong” from “the parser is wrong” in seconds instead of hours.

    Maximal munch, rule order and keyword ambiguity

    Maximal munch and rule order are the same problem looked at from two angles. A rule set that matches = before == produces two EQ tokens where one EQEQ was intended, and the parser downstream reports a confusing error two stages away from the cause. Sorting operator alternatives longest-first, and keywords before identifiers, eliminates both classes of bug without any lookahead machinery.

    Keyword ambiguity needs an explicit decision. Most languages treat keywords as reserved: the lexer recognizes if as a keyword in every position, so a variable named if is a syntax error rather than an identifier. Languages that allow contextual keywords push the decision into the parser instead. Either choice works; leaving it undecided gives you a lexer whose behavior depends on which rule happened to be listed first.

    Does tokenization change parser accuracy? Evidence from a 2026 study

    Tokenization quality isn’t cosmetic, and there’s now published evidence for that. Shamaeva and Loukachevitch (2026), in Pattern Recognition and Image Analysis, Vol. 36, pp. 486–497 (DOI 10.1134/S105466182670029X), compared two ways of evaluating a syntactic parser: with its built-in tokenizer, and with a tokenizer that returns gold markup. For a significant number of sentences, the built-in tokenization differed from the gold one, and average UAS and LAS scores were higher when the parser was evaluated with gold markup — some metrics by 0.05 or more.

    The study covers Russian-language corpora (SynTagRus, GSD, PUD, Taiga, and Poetry) and the parsers UDPipe, Stanza, Natasha, DeepPavlov, and spaCy, as they existed for this 2026 research. Those are the versions studied — not a claim about current releases of any of those tools.

    The practical takeaway for anyone building a parser: if your accuracy numbers look wrong, test the tokenizer before you rewrite the grammar. A meaningful share of apparent parsing errors originate one stage earlier, in segmentation decisions the parser never sees.

    Parsing Template Tokens and Placeholders Safely

    A document-numbering engine is a good concrete case. The tan-erp issue #6, opened 18 September 2026, asks how IDocumentNumberGenerator should parse template strings containing nine token forms — {PREFIX}, {BRANCH}, {YYYY}, {YY}, {BBBB}, {BB}, {MM}, {DD}, and {SEQ:X} — derive a period key from them, and generate unique document numbers atomically with full unit-test coverage.

    Four design decisions follow from that framing.

    Design around a registry, not a loose regex. A table of known token names with a resolution function per token makes unknown placeholders detectable. A single permissive pattern like \{[A-Z]+\} matches anything and tells you nothing about whether the result is meaningful.

    Anchor the pattern. Match whole placeholders with anchored boundaries, and handle escaping explicitly. If a substituted value can itself contain {, an unanchored second pass will re-parse it — that’s how injection-shaped bugs get into template engines.

    Fail loudly on unknown tokens. Raising is the correct behavior. Leaving the literal {BRANCH} in the output, or dropping it, silently produces malformed or colliding identifiers, and the failure surfaces much later in a way that’s hard to trace back to the template.

    Keep atomicity outside the parser. Parse first, then generate the counter inside a single transaction. The parser can be perfectly correct and two concurrent requests can still get the same number if the sequence increment isn’t atomic.

    There’s a harder adjacent case worth flagging: languages without whitespace word boundaries. The same issue pairs the template work with Thai tokenization, because a naive split on spaces produces no segmentation at all for Thai text. That branch needs a dictionary- or rule-based tokenizer — a different algorithm from placeholder substitution, and the two shouldn’t be combined into one pass.

    Unknown placeholders should raise, not fall through

    The difference between “raise” and “fall through” is the difference between a caught error and a data-integrity incident. A fall-through leaves the literal token in a generated document number, and every request that hits the same template produces the same literal — so the collision is systematic, not random. Raising converts invisible corruption into a stack trace at the first occurrence.

    Token Parsing in a Design-Token Build Pipeline

    A design-token parser reads generated platform artifacts — SCSS declarations, var() references, Kotlin token files — rather than a hand-maintained source file. That changes the risk profile: the input is machine-written, verbose, and changes shape whenever the upstream generator changes.

    The slint issue #3, opened 27 September 2026, documents the parsing contract for a Material 3 Expressive token pipeline. The source is Compose material3’s generated token files in androidx/androidx, pinned to commit 23327507f7fc7d5b19d65fec4b090f60c970079b (androidx-main, 2026-09-27), whose files carry the header // VERSION: 14_1_0. Since 2026-08-05 those files use inline value classes, so token values sit on get() = lines — a parser written for the older file shape silently reads nothing and reports no error.

    Three rules from that contract transfer to any design-token pipeline:

    1. Fail loudly on any token that cannot be parsed or mapped. Never silently drop or default a value. A dropped token becomes a component that renders with the wrong color, discovered by a designer rather than by CI.
    2. Reproduce committed output byte for byte in CI from the pinned commit. The generator runs on a clean checkout with one documented command, and the check passes only when the output is identical.
    3. Treat a version-pin bump as a deliberate, reviewed change. The pin is recorded in one place and printed into generated file headers, so bumping it is an explicit diff rather than an incidental drift.

    The referenced pipeline also resolves references between token files — a button token pointing at ShapeKeyTokens.CornerMedium, which points at ShapeTokens.CornerMedium — so the parser needs a typed intermediate model, not a flat name/value map.

    Why the pinned commit and the date are the parsing contract

    A design-token parser is only correct relative to a specific input. The same file path yields different values at different commits, and the file format itself changed on 2026-08-05. Recording 23327507f7fc7d5b19d65fec4b090f60c970079b and // VERSION: 14_1_0 in one place turns “the parser is broken” into “the pin moved and the diff is these lines.” Without the pin, a format change and a parser bug look identical from the outside.

    Build the Token Sets Once: Parsing Performance and Duplication Risk

    Token sets, DFAs, and grammar patterns should be constructed once and held on the parser instance. Rebuilding them at every call site is a performance problem, and duplicated parsers are a correctness problem that costs more than the performance one.

    The measured cost of getting the performance case wrong comes from the hefermotor issue #22, opened 19 September 2026. The _TokenSets primitive built an array literal on every method call, and the grammar called those methods at roughly 35 test sites. For the three largest sets — expr_start with 40 tokens, called at every call expression, plus case_pattern_start and type_start — building per call added 10–15% to the sum of Parse.tree over the standard library’s files in a debug build. Moving those three to fields on _Parser removed the overhead. The remaining, smaller sets were not measured.

    The correctness case comes from the plus-uno issue #622, from the 2026-09-18 architecture review. The docs token-literal check carried a parallel implementation of most of the tokens module: its own token reader, declaration regex, alias resolver with its own cycle guard, colour key, dimension key, family map, and SCSS brace parser. Removing the duplicate — tracked in PR #667 — affected 315 live token pairs that were equal before the change and unequal after. Unchanged output had been resting on no fallback literal pairing an opaque hex against a translucent token, not on any guarantee the code provided.

    That number is the argument. A duplicate parser isn’t just redundant; it’s a second set of semantics that diverges the moment either copy is edited, and the divergence stays invisible until someone enumerates it.

    Enumerate output changes instead of assuming they are absent

    The plus-uno acceptance criteria state the discipline directly: any output change is enumerated and reviewed, not assumed absent. That’s a higher bar than “tests pass,” because passing tests only prove the cases they cover. The 315-pair review is the model — the reviewer had to name the mechanism that preserved the output, not just observe that nothing appeared to break.

    The rule of thumb: one grammar module, one token registry, built once, version-pinned.

    How to Test a Token Parser: Acceptance Criteria from GitHub Tickets

    This section exists because of what the search results show about how the work is actually gated. Eight of the top ten results are GitHub issues and repositories, and five of them gate completion on tests passing or on enumerated acceptance criteria — but none presents those criteria as reusable guidance. The six criteria below are pulled from those tickets and generalized.

    1. A permanent regression test that fails without the fix and passes with it. The synchro issue #53 attaches this explicitly to the error-leak fix for audit finding C12, noting that no existing test covered the behavior. A test written after the fix that passes on the fixed code proves nothing on its own; it must be shown to fail on the unfixed code.

    2. Enumerate every output difference. The plus-uno criteria require that output changes be enumerated and reviewed rather than assumed absent. Claiming “no behavior change” without naming the mechanism is not evidence.

    3. Byte-for-byte reproduction in CI from a pinned upstream commit. The design-token pipeline’s reproducibility check is the strongest of the six because it admits no interpretation: either the generated bytes match or they do not.

    4. Fixture hygiene in the same change. When you tighten a predicate, update the success fixtures alongside it. The cortexfs fix added token_type: "Bearer" to the Codex OAuth success fixtures in the same change that required the field — otherwise the newly correct parser would have failed its own pre-existing happy-path test.

    5. Negative cases as first-class tests. Cover absent token_type, an unknown template placeholder, an unsupported token type such as mac, and a truncated Base64url segment. The cortexfs test plan adds a missing-token_type failure case to the existing hermetic OAuth regression; the synchro fix requires the defect to stop reproducing.

    6. A test for the error path itself. Assert that the client-facing error is bounded and contains no parser internals. This is the criterion that would have caught C12 before an audit did.

    Negative tests: absent fields, unknown tokens, truncated segments

    Negative cases deserve their own emphasis, because they’re where the parser’s boundaries actually live. Positive tests describe what a parser accepts; negative tests describe what it refuses, and refusal is the behavior that protects everything downstream.

    Four cases cover most of the risk surface. An absent required field — the token_type case. An unrecognized token in a registry-driven parser — the template placeholder case. A recognized-but-unsupported token — the mac case, where the field is present and well-formed but the implementation doesn’t support the semantics. And a truncated input — a Base64url segment whose padding has been stripped or whose length is wrong. Each should have a named test, and each test should assert on the specific error the caller receives, not just that an exception was raised.

    Conclusion

    A token parser isn’t one thing. It’s four different parsers sharing a name, and the format you’re decoding decides which one you need, what its output may claim, and where it has to stop.

    Name your token format first, then use the decision table to pick the branch. If it’s a JWT or OAuth token, decode it for inspection but verify the signature before you trust it, and reject a missing or unrecognized token_type instead of defaulting to Bearer. If it’s a lexical, template, or design token, define the token registry, build it once, fail loudly on anything unrecognized, and back the result with a regression test that fails without your fix.

    FAQ

    Is a token parser the same thing as a tokenizer or a lexer?

    They overlap, but they aren’t identical. A tokenizer or lexer produces the token stream; a token parser consumes tokens and gives them structural meaning. In practice, plenty of codebases use the terms interchangeably, so read the token format before assuming which one a given library means.

    Does a JWT parser verify the token’s signature?

    By default, no. Decoding reverses Base64url and deserializes JSON — nothing more. Verification requires the HMAC secret or RSA public key, and a decoded token can carry any claims its sender invented. For anything security-relevant, use libraries with signature verification explicitly enabled — the same decode-only boundary applies to any online JWT parser you use for inspection.

    Why does RFC 6749 require token_type, and is it case-sensitive?

    RFC 6749 §5.1 requires both access_token and token_type in a successful response so the client knows how to present the token. token_type is case-insensitive: Bearer, bearer, and BEARER are all valid. RFC 6749 §7.1 adds that a client must not use an access token whose type it does not understand — so reject rather than default.

    Why shouldn’t a token parser’s error message include its internal token details?

    Parser internals leak grammar rules, offsets, and token names, which hands an attacker a map of the input format. The synchro issue #53 records an audit-confirmed finding — C12, verdict CONFIRMED, severity Low — of exactly this kind of leak going out to clients. Return a bounded generic error and keep detailed diagnostics in server-side logs behind a stable correlation ID.

    Can I parse template placeholders like {BRANCH} or {SEQ:X} safely, and what about ones I don’t recognise?

    Parse against a registry of known tokens rather than a loose regex, and anchor the pattern. Raise on any unrecognized placeholder instead of leaving it in the output — silent passthrough produces duplicate or malformed identifiers. Handle escaping explicitly so injected values containing braces aren’t re-parsed on a later pass.

    Should I write my own token parser or use a library?

    Use a library when a standard defines the format — JWT and OAuth have mature, audited implementations, and hand-rolling the security path is where bugs live. Write your own when the token set is small, the grammar is yours (template placeholders, a domain-specific syntax), or you need it to fail loudly on unknown tokens in a way no general library does. Either way, build token sets once and cover the negative cases with tests.