Master the fundamental concepts of lexical analysis through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksReal compilers encounter garbage characters, unterminated strings, and malformed numeric literals constantly. Clang's lexer reports precise diagnostics and recovers so parsing can continue. A lexer that aborts on the first unexpected byte makes the compiler useless on incomplete or generated code.
When the lexer sees an invalid character, it should record an error with line and column, skip or consume the offending input, and return either an error token or continue scanning. Never read past the buffer end. Pair every error with enough context that the user can fix the source.
cLoading…
TOK_ERROR or a dedicated invalid token for the parser to handleProduction compilers embed this step inside a longer pipeline. GCC flows through cpp, cc1, assembly, and ld; Clang uses the driver, Sema, LLVM IR passes, and a target backend. LLVM bitcode, JVM bytecode, and WASM are other familiar IRs at the same layer. The exercise isolates one pass so you can test it alone before chaining it to the next stage.
You will add error detection and reporting to your lexer so invalid input produces clear messages instead of silent mis-tokenization. This exercise requires handling unexpected characters and malformed literals while keeping the scan loop running.
Extend the lexer from the previous task so it reports errors and keeps going. A real compiler shows every lexical mistake in a file, not just the first one.
A source file on stdin, using the same token set as the previous task (keywords int return if else while, operators + - * / = == != < > <= >= && ||, delimiters ( ) { } ; ,, numbers, identifiers, // comments), plus two additions:
"...". A backslash escapes the next character, so "say \"hi\"" is one string. A string may not contain a raw newline./* ... */, which may span lines.| Error | Trigger | Message | Reported at | Where scanning resumes |
|---|---|---|---|---|
| Invalid character | a character that starts no token (@, #, $, a lone !, & or |) | invalid character 'X' | that character | the next character |
| Unterminated string | " with no closing " before the end of the line or file | unterminated string | the opening " | the newline (or EOF) |
| Unterminated comment | /* with no closing */ | unterminated comment | the / of /* | EOF |
| Invalid number | a digit run immediately followed by letters, digits or _ | invalid number 'TEXT' (whole run, e.g. 123abc) | its first digit | after the run |
Positions follow the previous task: lines and columns start at 1, a newline goes to the next line and column 1, every other character is one column.
One line per error, in the order they occur:
cLoading…
Then a summary line:
cLoading…
N counts the valid tokens (strings included, errors and EOF excluded). M is the number of error lines printed. A clean file prints only the summary.
Input:
cLoading…
Output:
cLoading…
The 10 valid tokens are int x = ; on line 1, char c = on line 2 (char is just an identifier here), and return x ; on line 3.
LexerError record (message, line, column) collected into an array as you scan; print them after lexing finishes.Hidden tests put several errors on one line, use _ inside a number, and leave a block comment open until EOF.