Master the fundamental concepts of cpython internals through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksBuiltins like len, abs, and sum are C functions registered on the builtins module. They parse PyObject* arguments, do native work, and return new Python objects without ever entering the bytecode interpreter loop.
A builtin uses the same C API as extensions but lives in Python/bltinmodule.c or similar:
cLoading…
Every builtin you study follows the same checklist:
PyArg_ParseTuple with a format string like "i" or "O"PyLong_Check, PySequence_Check, or similar type guardsPyLong_FromLong) or NULL with PyErr_SetStringType checks, error handling, and correct refcounting are mandatory. Fast paths often bypass generic PyNumber dispatch when the argument type is known.
For example, abs(-42) calls C code that reads the PyLong digits directly instead of re-entering the interpreter loop.
Fast paths should validate types once, then read internal struct fields directly. Falling back to abstract PyNumber APIs preserves semantics but costs the speed that motivated writing C in the first place.
This exercise asks you to implement a Python builtin in C following CPython conventions. You will write argument parsing, type validation, and return value construction for a function callable from the REPL.
You will use the same mental model here when reading production interpreter source later in the track. Sketch one concrete input on paper, predict the outcome, then confirm with code. That discipline catches logic errors early and makes debugging far faster when you extend the implementation in follow-on tasks.
Implement a new built-in, reverse(s), the way Python/bltinmodule.c implements builtins, together with the few builtins needed to inspect strings: len, ord, chr, getsizeof and kind. The real work is Unicode. A CPython str is a sequence of code points, not bytes. PEP 393 stores it with 1, 2 or 4 bytes per code point, chosen by its largest character. So reverse must decode the UTF-8 source, reverse the code points, and the result keeps the same storage kind.
One expression per line: a string literal, an integer, or a call name(arg, ...) whose arguments are expressions (calls nest).
'...' or "...". They contain UTF-8 text and the escapes \\n \\t \\\\ \\' \\" \\xNN \\uNNNN \\UNNNNNNNN.-.| call | result |
|---|---|
reverse(s) | a new str with the code points of s in reverse order |
len(s) | the number of code points |
ord(c) / chr(i) | code point ↔ one-character string |
getsizeof(s) | sys.getsizeof on 64-bit CPython 3.8+: ascii 49 + n, latin1 73 + n, ucs2 74 + 2n, ucs4 76 + 4n |
kind(s) | 'ascii' (max < 0x80), 'latin1' (< 0x100), 'ucs2' (< 0x10000) or 'ucs4' |
The empty string is ascii.
TypeError: NAME() takes exactly one argument (N given)TypeError: reverse() argument must be str, not int (the same for getsizeof and kind)TypeError: object of type 'int' has no len()TypeError: ord() expected a character, but string of length N foundTypeError: ord() expected string of length 1, but int foundTypeError: an integer is required (got type str) (from chr)ValueError: chr() arg not in range(0x110000)NameError: name 'X' is not definedSyntaxError: EOL while scanning string literalSyntaxError: invalid UTF-8 in source (malformed, overlong or surrogate encodings)SyntaxError: illegal Unicode character (\\U above 0x10FFFF)SyntaxError: truncated \\xXX escape (likewise \\uXXXX, \\UXXXXXXXX) when the escape lacks hex digitsSyntaxError: unsupported escapeSyntaxError: invalid syntaxPrint each result's repr:
' and no ". The chosen quote and \\ are escaped with a backslash.\\n \\t \\r stay escapes. Other characters below 0x20, and 0x7F..0x9F, become \\xNN, and surrogates (from chr) become \\udNNN, all in lowercase hex.Input:
cLoading…
Output:
cLoading…
PyUnicode_New(size, maxchar) does.Hidden tests cover 4-byte characters (emoji), a combining accent that moves when reversed, escapes and quote choice in repr, a lone surrogate from chr, and the type, value, name and syntax errors.