Writing Documentation¶
scikit-agent follows scikit-learn’s guidelines for contributing documentation, which means that if you have written documentation for a scientific Python project before, most of what you know already applies. This page records the parts that are specific to this project: the conventions our toolchain forces, the vocabulary the documentation shares, and the register each kind of page is written in. Where this page and the scikit-learn guidelines disagree, this page is the one to follow, and the last section says where the disagreements are and why each one exists.
You do not need to read it through before contributing. It is reference material, and the sections stand on their own.
The four surfaces¶
Documentation lives in four places, and which one you are editing decides most of the questions below.
Surface |
Where it lives |
What it is |
|---|---|---|
API documentation |
docstrings in |
the contract of a module, class, method or function |
User guide |
|
how part of the library is used, and why |
Examples |
|
runnable programs, rendered into the gallery |
Other documents |
|
project documents and the API index |
There is a fifth directory that looks like a surface and is not.
docs/auto_examples/ is generated by Sphinx-Gallery from examples/ on every
build, it is not in the repository, and anything written there is thrown away by
the next build. If you want to change what a gallery page says, change the
example it is generated from.
Voice¶
Three of those surfaces have different readers, so they are written differently. The register follows the reader, which means the gallery and the user guide are not supposed to sound alike.
The gallery teaches, and it reads like a well-written undergraduate textbook. Its reader is capable and new to the material: they have calculus and some probability, they have not used this library, and they have not read the papers behind it. So a gallery page motivates a thing before it formalizes it, defines a term where the term first appears, and carries its worked example far enough that the reader sees the numbers come out. It builds in order, using only what its earlier sections have established.
Three neighbouring registers are worth naming, because writing in one of them by accident is easy. A gallery page is not a paper, so it does not position itself against a literature or argue for its own novelty. It is not a blog post, so it does not make jokes, asides, or appeals to how interesting the material is. It is not a reference manual, so it explains in prose instead of enumerating.
The user guide is written for a developer rather than a student. Its reader knows Python and their own field, has a task in mind, and wants to know what the library calls a thing, what to call, what it does and what it refuses. A user guide page therefore gets to the code early and leaves the derivation out. When a reader wants the model itself rather than the API around it, the page says so and sends them to the gallery.
The community pages, including this one, are written for someone deciding how to contribute and what the project will expect of their work. They state a convention and the reason for it, and they are read once rather than kept open.
Some things hold across all three. Write plainly, in complete sentences, and make each claim once, in the order a reader needs it. Say what a thing does before you say why it matters. Use “we” for a step the page takes together with the reader, “you” for something the reader does in their own code, and avoid the passive where it would hide who is acting.
Three habits work against all of that, and they are worth watching for in docstrings as much as in pages.
The first is the slogan: a bolded, compressed assertion at the head of a paragraph that the reader has to unpack before they can read on. Bold is useful for naming a thing — a class, a mode, a parameter — and unhelpful for carrying a proposition. Open with an ordinary sentence saying what the paragraph is about.
The second is the fragment. Body text is made of complete sentences; headings, table cells and short labels may be noun phrases. A line that drops its subject or its verb to sound terse is usually a sentence worth writing out.
The third is emphasis doing the work an explanation should do. Capitals on a claim, and the “not X, but Y” reversal used for effect rather than for a real contrast, both assert where the reader wanted to be shown something.
Two markup languages¶
Pages under docs/ are MyST Markdown. Docstrings are reStructuredText, read by
sphinx.ext.napoleon. Most of the time the difference does not come up, but it
decides how you spell a cross-reference and how you spell mathematics:
In a page |
In a docstring |
|
|---|---|---|
Cross-reference |
|
|
Mathematics |
|
|
Docstrings¶
A docstring is a contract. It says what a caller may rely on, which is a different job from explaining why the code changed or walking a reader through a particular model. Those belong in the pull request, the user guide or an example.
The format is numpydoc, and the section order is numpydoc’s: Parameters,
Returns, Raises, See Also, Notes, References, Examples. The pair most often
written backwards is the last two of those middle ones — See Also comes before
References — and some docstrings in the tree still have it the other way
round.
Parameters and attributes are written as name : type, default=value, with the
description indented underneath:
sample_count : int, default=1000
The number of independent paths the simulator draws.
method : {"exact", "neural", "tabular"}, default="exact"
Which best-response method solves each decision.
state_grid : Grid or AxisSpec
The states the backup is taken over.
values : ndarray of shape (sample_count,) + entity_shape
One value per path and per entity instance.
decision_rules : dict of str to callable, default=None
A rule per control symbol. None means every control is solved for.
The type is spelled by a short set of conventions, which exist so that two docstrings describing the same thing describe it the same way:
Python’s own type names, so
boolrather thanbooleanandstrrather thanstring.Shapes in parentheses after the type, with the axes named the way the library names them:
sample_countfor the sample axis andentity_shapefor the trailing per-instance axes, followingSimulator.A fixed set of string options in braces, as
{"exact", "neural"}.ofas the delimiter for containers, solist of stranddict of str to callable.ndarray,Tensororxarray.DataArraywhere only one of them will do, andarray-likewhere any of them is accepted.Library types under their class name, which also makes them cross-reference:
GroundedBlock,Grid,Control.default=Nonelast, with the description saying whatNonemeans. A default that is notNoneis written as it would be typed.
A See Also section gives one reference per line, each with a colon and a short
reason for the reader to follow it:
See Also
--------
solve_in_order : Solve named decisions in a given order.
solve_symmetric_equilibrium : Solve one decision against copies of itself.
An Examples section is optional, and when there is one it has to run as
written, imports included, because it can be executed:
pytest --doctest-modules src/skagent/rule.py
If an example cannot run without a model built over several lines, that is a sign the explanation wants the user guide or the gallery, where the model can be built in the open. A fragment that would fail the command above is better deleted than left in place.
Docstrings are code, so ruff wraps them at 88 columns rather than 80.
User guide pages¶
A user guide page opens with what the thing does and when a developer would reach for it. A shape that works, and that most of the existing pages follow:
A short explanation of what it does, in the library’s own vocabulary.
What it is good for, and when to reach for the alternative instead.
The code: how it is called, with one or two short examples that paste.
What it refuses, and what it costs — a complexity, or a rule of thumb about problem size, where one is known.
A figure, where one helps, drawn from a shipped model rather than invented for the page.
A pointer to the gallery page that derives the model or the method. The derivation belongs there; a user guide page cites mathematics rather than developing it.
Prettier wraps markdown prose at 80 columns with --prose-wrap=always, so write
the paragraph and let the hook lay it out rather than hand-wrapping against it.
Inline code takes single backticks.
Introduce a term on the page that owns it, and use it as defined everywhere else. The solver vocabulary is the worked case: method and schedule are introduced in Solvers Guide, and a page about something else uses them as defined there instead of redefining them in passing. For the objects in the terminology table below, use the definition in the table.
Material that only some readers want can be folded into a dropdown:
:::{dropdown} Where the bound comes from
The derivation.
:::
Dropdowns suit references, derivations and use-case narrative. They come with two limits. A page’s examples stay visible and are never folded, since they are what most readers came for. And a cross-reference target inside a dropdown does not resolve, so anything that refers to it has to be folded along with it.
Examples¶
An example is a runnable Python file in examples/, with its prose in the
module docstring and in # %% cell comments. A file named plot_*.py is
executed when the documentation is built and its figures are captured; anything
else is rendered without being run.
That execution makes an example a test of the prose around it. A number quoted in the text is a number the reader will see printed, so a change that moves the number and not the sentence will be visible on the published page.
A gallery page is read start to finish by someone meeting the model for the first time, which is why the textbook register belongs to this surface and not to the user guide. Such a page states the problem, says why it is hard, shows what the library does about it, and ends with what the numbers show. A reader should be able to follow it without opening the source of anything it calls.
References¶
Where a work has an identifier, cite it with the Sphinx role, in pages and docstrings alike:
:arxiv:`2106.03958`
:doi:`10.1016/j.jmoneco.2021.07.004`
A work the library as a whole builds on belongs in the bibliography in
Ecosystem. A work behind one function belongs in that function’s
References section.
Cross-references¶
Objects, pages and sections are each referred to differently:
Target |
How to write it |
|---|---|
Object |
|
Page |
|
Section |
|
The leading ~ on an object reference prints the last component only, which is
usually what you want in a sentence. The explicit py: form is the majority
spelling in the tree and cannot be confused with a MyST extension role, so
prefer it to the bare {class}.
One thing not to do: renaming an existing target or page. Other pages and external links point at them, and a rename breaks both without any warning.
Building and checking¶
uv sync --extra docs
uv run --no-sync python -m sphinx -b html -W --keep-going docs docs/_build
uv run --no-sync python -m http.server 8000 -d docs/_build
That is the command CI runs, writing where CI writes, and make html and
make livehtml in docs/ build into the same place. If you point Sphinx
somewhere else, remember that the page you open has to come from the build you
just ran; a stale index.html left in another directory looks exactly like a
change that did not take.
-W turns warnings into errors, which is what CI does, so a broken {doc} or
{ref} link fails the build instead of shipping. References to Python objects
are the exception, and worth knowing about: one that does not resolve renders as
plain code and says nothing, so check that a new one is a link in the built
page.
The first build executes every gallery example and takes minutes. Later builds reuse the cache and take seconds, and the cache is dropped automatically when the code that draws a model diagram changes.
Before opening a pull request that touches documentation, check that the build
passes with no new warnings, that you have looked at the rendered page rather
than only the source, and that any docstring example you added or changed passes
pytest --doctest-modules on its module. Prefix the pull request title with
DOC when it changes documentation only.
Terminology¶
These are the words the documentation uses for the library’s objects, with one definition each. Where a page needs one of them, this is the meaning to use.
Term |
Definition |
|---|---|
block |
A set of variables and the equations relating them. |
|
One period of model behavior: its shocks, dynamics, controls and rewards. |
|
Several blocks chained in order, so that one’s outputs are the next one’s inputs. |
shock |
A variable drawn from a distribution rather than computed. |
dynamics |
The equations of a block, read in declaration order. |
control |
A variable an agent chooses, declared with an information set instead of an equation. |
information set |
The variables an agent may read when choosing a control. |
reward |
A variable an agent is paid in. An agent’s payoff is the sum of the rewards it owns. |
agent |
A role that owns controls and rewards. A role is not a population. |
arrival state |
A variable read before it is assigned, carrying a value in from the previous block or period. |
calibration |
The parameter values a block is read against. |
|
A block together with the calibration and shock generator it is read against. |
|
A grounded block plus a discount variable: one period of a dynamic program. |
decision rule |
A function from an information set to a control’s value. |
policy profile |
One decision rule per control in the model. |
|
A labeled, discretized set of points a model is solved or evaluated over. |
method |
An object that solves one decision, given rules for the others. |
schedule |
A function that decides which decision is solved next, and when to stop. |
entity class |
A declared population; a variable in a block carrying one is an attribute of each instance. |
crossing |
An equation reading out of an entity class into a variable that has no such axis. |
Where this differs from scikit-learn¶
Each of the differences below is forced by this project’s toolchain or its size, rather than chosen as a preference.
Our pages are MyST Markdown and our docstrings are reStructuredText, so
scikit-learn’s page mechanics — .. dropdown::, :ref:, .. currentmodule:: —
appear here in their MyST spellings.
They ask for 88 columns throughout. Here prettier owns markdown and wraps it at 80, while ruff owns code and wraps it at 88, so the number depends on which file you are in and neither is a judgment call.
numpydoc itself is not installed. Napoleon reads the same sections, so the format is the same one, but nothing validates a docstring against it. The section order and the parameter format are conventions here, upheld by review.
There is no glossary yet. The terminology table above is the definition list,
and {term} does not resolve. Promoting the table to a glossary page is worth
doing once the vocabulary settles.
Changes here are not approved by two core developers. How pull requests are reviewed is in the Community Guide contributor guidelines.
The teaching lives somewhere else. scikit-learn’s user guide teaches the algorithm, which is why their content order puts the mathematics into it. Here the gallery does the teaching and the user guide is developer-facing, so the two have different content orders and only the gallery carries a derivation.
Finally, this standard binds docstrings written or edited from now on. Bringing the existing ones into line with it is separate work.