/take-notes
Turn a source into
notes you can learn from — not a transcript dump, not a
one-paragraph summary. The output is one self-contained HTML page written to
~/take-notes/html_reports/
and opened in the browser, so the notes accumulate
into a browsable local archive instead of scrolling away in the terminal.
Invocation:
/take-notes <url> [more urls…] [focus]
. If no URL is given, ask for
one. Several URLs are
one note about one subject from several sources — a
talk and the deck it was given from, a paper and the repo that implements it —
not one note each; Step 1 says how they combine.
The optional focus does two things: it narrows what Step 1 asks the source for
— on a long or multi-topic source, that's the difference between fetching the
whole thing and fetching only the part that matters — and it narrows what the
finished notes emphasise in Step 3-4. Skip it to cover a source in full; add it
("just the API design part") when only part of a long source is relevant.
,
, and
manage the tag vocabulary instead —
see Step 0.
Resolve (before any command, both source types)
The scripts are bundled with this skill, a direct sibling of this file. Set
to the
absolute path of the directory containing THIS SKILL.md you
just Read — your harness reported it in the Read result — and substitute it
literally in every command below:
bash
SKILL_DIR="<absolute path of the directory containing the SKILL.md you Read>"
if [ ! -f "$SKILL_DIR/scripts/render.py" ]; then
echo "ERROR: scripts/render.py not found under SKILL_DIR=$SKILL_DIR" >&2
exit 1
fi
Step 0 — tag management short-circuits everything else
Three invocations manage the tag vocabulary instead of writing a note. If the
invocation is one of them, run the matching command, report the result, and
stop — no source, no note, nothing else in this file applies:
| Invocation | Command |
|---|
| uv run "${SKILL_DIR}/scripts/tags.py"
|
/take-notes --add-tag "AI"
| uv run "${SKILL_DIR}/scripts/tags.py" --add "AI"
|
/take-notes --remove-tag "AI"
| uv run "${SKILL_DIR}/scripts/tags.py" --remove "AI"
|
| re-files existing notes — the multi-step pass below |
Both editing forms are repeatable — pass
or
once per tag.
The script prints the resulting vocabulary; report that, and nothing more. It
rewrites only the
key, so
survives untouched.
cannot be removed: it is the fallback the note writer needs when a
source fits nothing. The script says so and leaves it in place.
— re-file the notes already on disk
Filing a note under a new tag used to mean re-running
on its
source: a refetch and a full rewrite, to change one word in the rail. This pass
edits the rendered notes instead. Run it after adding tags to a vocabulary that
was empty or thinner when those notes were written.
- Read the vocabulary —
uv run "${SKILL_DIR}/scripts/tags.py"
. If the
only entry is , say so and stop: there is nothing to file notes
under yet, and the user needs first.
- List what is on disk —
uv run "${SKILL_DIR}/scripts/retag.py" --list
.
One JSON object per note: path, title, byline, kind, date, current ,
an , and .
- Choose from the vocabulary and nothing else. For every note with
, pick the entry that fits from the list read in step 1;
add further tags after the primary when they genuinely apply. The list is
closed — the script rejects anything not on it rather than inventing a tag
that would exist on one note and in no chip. Nothing fits, leave it on
; a wrong file is worse than an unfiled note.
- Write each one —
uv run "${SKILL_DIR}/scripts/retag.py" --set "<path>" --tag "<primary>" [--tag "<extra>"]
- Rebuild the gallery so the chips match the notes —
uv run "${SKILL_DIR}/scripts/gallery.py"
- Report one line per note re-filed, plus how many were left on .
means the note already carries a deliberate tag. Leave
those alone unless the user asked for every note; overwriting a filing someone
chose is not an update.
The pass rewrites only the rail's tag row. No source is fetched and no prose is
regenerated, so it costs the listing and the model's choices, nothing more.
Step 1 — route to the right acquisition guide
Pick one reference per source by looking at it, Read it, and follow it. Only
the acquisition differs; everything after Step 2 is identical for every source.
Match
top to bottom and stop at the first row that fits — arXiv, Slides and
GitHub links are
pages too, so the catch-all row would swallow them.
| Source | Read |
|---|
| YouTube URL, any other video URL yt-dlp supports, or a local media file | |
| (or an / arXiv DOI link) — a paper | |
docs.google.com/presentation/...
— a slide deck | |
github.com/<owner>/<repo>
— a repository root, not a file, PR, or issue | |
| Any other page — blog post, docs page, news article | |
Each guide hands back the same thing, and nothing more:
- title
- byline — channel for video, author or site for an article, the paper's
authors, the repo's owner
- span — duration for video, publication date for an article, submission
date for a paper, latest release for a repo
- canonical URL (plus the YouTube video ID when there is one)
- body — the timestamped transcript, the article text, the paper full text,
the deck's slides and speaker notes, or the README plus the repo's structure
Video sources also hand back, when yt-dlp reports them: channel URL,
published date, views, a thumbnail URL, and the caption language.
Pass these to Step 5 too — they drive the two-pane video layout. Articles never
have them; leave those flags off entirely rather than passing empty strings.
Captions arrive in the language actually spoken. When the guide reports the track
was machine-translated (no original-language track existed), note it in
Going deeper: translated captions mangle proper nouns, so names taken from them
are unreliable and quotes are twice-removed from what was said.
If a guide reports it could not get the body, say so and stop. Never write
notes from a title, a description, or a paywall stub.
More than one source
Route each URL through its own row above and collect the same field set for
each. Fetching is the only part that repeats: from Step 2 on there is one
language, one tag, one body, one note.
The first URL is the primary source. Everything the note's chrome shows
comes from it — title, byline, span, canonical URL — and it picks the layout: a
video first renders the two-pane video note with a poster and timestamps, a
deck, paper or article first renders the article one. The rest are
companions: they contribute body, and Step 5 links them in the rail. Order
is the user's control over that, so take it literally rather than promoting the
richest source.
A companion that yields no body is not fatal: name it, say the notes are
poorer for it, and write from what did arrive. A primary that yields no body
stops the run — the note would be filed under a source it was not written from.
A focus narrows every source at once, which is where it earns the most: two
full-length sources is the largest input this skill ever takes.
Do not put note-writing guidance in the reference files, and do not put
acquisition detail here. Two copies of the writing standard will drift.
Step 2 — settle the language and the tag
One read of
answers both:
and
.
Language
This skill writes in English or Spanish only. Resolve which, in this order,
and stop at the first that applies:
- or in the invocation. Wins over everything,
including a config set to — an explicit flag is not a question.
- A language named in plain words in the invocation ("take notes on this in
English"). Same standing as the flag; if somehow both appear, the flag wins.
- . Read it. or → use
it. → go to the question below.
- Default: English. No config, an unreadable one, or any other value — a
broken config must never block the run.
json
{ "language": "es", "tags": ["Unknown", "AI", "Investing", "Engineering"] }
with anything other than
or
is
not an error to stop on,
and must not be passed through:
silently falls back to English
chrome for unknown codes, which would pair English furniture with prose in a
third language. Say the value is unsupported, resolve from step 3 onward, and
name what you used instead:
is not supported (English or Spanish only) — writing in Spanish per your config.
When you resolve to a language
without asking, say so in one short line.
Point at the config file only when the language came from the config or the
default — someone who just typed
does not need to be told how to
set a preference they have overridden:
Writing in English (default). Set
in
to change.
If the source is not in the language you resolved to, say that too — a
Spanish video silently producing English notes is the one surprise worth calling
out:
Source is in Spanish; writing in English per your config.
When the config says
Ask once, with
, before writing anything. Offer exactly two
options — English and Spanish — nothing else. Put the source's own language
first, labelled "(Recommended)" (e.g. a Spanish-language video →
before
); if the source is in neither, put
English first.
However it resolves, the result sets the language for Step 4's headings and
prose, and the
code for Step 5 (
or
).
Tags
in the same file is a
closed vocabulary, curated by hand. Pick from
it; do not extend it:
- One primary tag — the single best fit for what this source is about.
That is what the gallery card shows and files the note under.
- Optional extras, only when they genuinely apply. Two is usually plenty;
tagging a note with half the vocabulary makes every filter useless.
- Never invent a tag. A name that is not in the list is not an option, no
matter how well it fits.
- Nothing fits, or is absent, empty, or unreadable → .
Silently. Do not ask, do not suggest a new tag, do not explain the fallback.
Say which primary tag you chose in the same short line as the language, without
justifying it: Writing in English (default), filed under Engineering.
Step 3 — read for teaching, not for summarising
Before writing, decide: what does someone who consumed this source now know
that they didn't before? That answer is the takeaway, and everything else
supports it. Note where the source explains a mechanism (goes in How it works),
defines jargon (Concepts), or leaves something unresolved (Going deeper).
With several sources, read them against each other before writing — that
comparison is the whole reason they were combined:
- Overlap — write it once, from whichever source explains it better. A deck
bullet and the sentence spoken over it are one point, not two.
- Gaps — a figure that is on a slide and in no transcript, a number said out
loud that is on no slide. These are what the second source bought.
- Contradictions — say so and attribute both. A talk that updates its own
deck is worth a line in Going deeper.
Never organise the notes by source. One set of sections, ordered by what has to
be understood first; a reader should not be able to tell where the seam was.
Step 4 — write the notes as HTML
Use the sections below. Write
body HTML only — no
,
,
, no
, and no metadata line: the renderer supplies the document
shell and the masthead from the fields you collected in Step 1.
There is no Markdown step. Emit the tags directly; nothing parses Markdown here,
which is why this skill needs no conversion dependency.
Step 5 — render it
Pipe the body HTML to the renderer, filling the flags from your Step 1 fields:
bash
uv run "${SKILL_DIR}/scripts/render.py" \
--title "<title>" --byline "<channel or author>" \
--span "<duration or publication date>" --url "<canonical URL>" \
--tag "<primary tag>" <<'HTML'
<h2>Executive summary</h2>
...
HTML
Pass one
per tag chosen in Step 2,
primary first —
--tag AI --tag Engineering
. With no
at all the note is filed under
.
The masthead flags describe the
primary source. When the run combined
several, add one
--source "<label>" "<url>"
per companion, in the order they
were given:
bash
--source "Slides" "https://docs.google.com/presentation/d/<DECK_ID>/edit"
They render as a short muted list under the source link. The label names the
kind of source —
,
,
,
,
— in the
note's own language; the title is already the
, and repeating it there
tells the reader nothing. A companion the run failed to fetch gets no
entry: the rail lists what the notes were written from.
For video sources, also pass whichever of
,
,
,
,
,
the guide reported.
is what switches the
rail to a poster + index; without it, the same two-pane layout renders for
articles instead, with a byline kicker and a numbered index in place of the
poster and timestamps.
For videos, pass
raw values and let the renderer localise them:
(seconds),
,
. It writes
11 min · 16 ago 2026 · 13.2K visualizaciones
for
and
11 min · Aug 16, 2026 · 13.2K views
for
.
stays a free-form
string for articles, whose span is a publication date rather than a length.
It writes
~/take-notes/html_reports/YYYY-MM-DD-<slug>.html
and opens it.
Re-running on the same source the same day
updates that file rather than
adding a near-duplicate; the script prints
or
with the
path. Report that path.
That is also how a note gets re-tagged: while the body is still in context,
re-run this command with a different
. Rewriting the tag inside an
already-written file is not something this skill does — re-run the source.
Pass
matching Step 2's choice (
or
). Add
to skip
the browser,
to write somewhere other than
~/take-notes/html_reports
.
Sections
Mandatory, in this order. The title and metadata line are not in the body —
they come from the renderer flags.
-
<h2>Executive summary</h2>
— 3–5 sentences: what the source covers and what
it argues.
-
<h2>The one takeaway</h2>
— 1–2 sentences wrapped in
. The single
most important insight. If you can't name one, the notes aren't ready.
-
— a
of 5–10 items, each
<li><strong>Claim</strong> — the detail that supports it</li>
.
Cap at 10; more than that is a transcript with bullets in front of it.
-
The outline, rendered to match the source:
- video →
<h2>Timestamped outline</h2>
, one per topic:
<li><a href="https://youtu.be/<ID>?t=754s">12:34</a> — <strong>Topic</strong> — one-line summary</li>
Use absolute URLs so the links jump to the right moment.
- article → , one per section:
<li><strong>Section heading</strong> — one-line summary</li>
, wrapping the
heading in when the page has stable anchors.
Aim for 6–15 entries either way; group adjacent material covering one idea.
The outline follows the
primary source only — it is one source's spine,
and interleaving two makes it navigate neither. A companion stays traceable
through inline deep links wherever a point comes from it: a slide's
<a href="<deck URL>#slide=id.<PAGE_ID>">
, a video's
.
Optional — include only when the source actually earns it, never as an empty heading:
- — jargon the source assumes or introduces, as
<li><strong>term</strong> — definition</li>
. Include a term only if not
knowing it blocks understanding the notes.
- — an for a mechanism, pipeline, or worked example
the source demonstrates. Code goes in .
- — what the source leaves open: unanswered questions,
claims made without evidence, and the concrete next thing to read or try.
Source figures —
and
return the diagrams, charts, and
screenshots the page carried;
returns an image URL for every slide.
Include one only when it is load-bearing — the diagram
is the explanation, the
chart
is the evidence — never a decorative photo, a header banner, an author
headshot, or (for a deck) a slide that is just bullets you already wrote out.
Cap at 3, the same "more than that is a dump" discipline as Key Points. Not a
section of its own: place
<figure><img src="<url>" alt="<alt text>"><figcaption>caption</figcaption></figure>
inline, in whichever section it supports — most often
How it works,
Key
points, or
Concepts. Each guide says how to confirm the URL really serves an
image before you embed it; a broken-image icon teaches nothing.
Rules
- Didactic means explaining, not compressing. A bullet only someone who already
consumed the source would understand has failed. Expand the reference; don't
preserve the author's shorthand.
- Learner's order, not source order. Only the outline follows the source's
sequence. Everything else is ordered by what has to be understood first.
- Quote sparingly — one or two lines that lose meaning when paraphrased.
- Own the notes. No "the speaker says that…" throughout; state the content and
attribute only genuinely contested claims.
- No padding. No "In conclusion", no restating the summary at the end, no bullet
whose content is "this is important".
- Flag the source's limits when it asserts things without support — that belongs in
Going deeper, and it's the part that makes the notes worth keeping.
- Language: write headings and body in whichever of English or Spanish was chosen
in Step 2; the structure doesn't change. Pass the matching ( or )
to the renderer.
- Keep the HTML plain: headings, paragraphs, lists, , , links,
, , simple tables, and (every source but video)
<figure><img><figcaption>
for a source figure. No inline attributes, no
, no classes — the stylesheet already handles presentation, and a note
that fights it will look wrong in dark mode.
- Escape what you write: , and must be , , —
in prose, in code samples, and in attribute values like an URL
(image URLs routinely contain an unescaped in their query string). The
renderer escapes the masthead fields but passes the body through untouched.
Related
- — frames and transcript. Use it directly when the question is visual.
includes its own copy of the transcript path so it runs standalone;
neither skill depends on the other being installed.
- — files a short webpage summary straight into the
personal Notion database via MCP. is the long form and stays local:
a full study page in
~/take-notes/html_reports/
, reviewed and edited before
anything is worth filing.