← volod.org

Getting My Slack Out of Slack

I wanted my AI to know what happened at work yesterday, and most of what happens at work happens in Slack. Decisions, half-decisions, "let's not ship this yet". So Slack was the source I cared about most, and it was the one I couldn't get out in a way I liked.

There was already a Slack connector in Vana. It runs slackdump, a good open source tool, and you give it a token copied out of your browser. That works. I just didn't want to paste a session token for my work Slack into anything, and I knew I'd never keep doing it every time the token expired.

Slack's own export doesn't help either. It's built for workspace admins, and what I wanted was my own view: my DMs, my threads, the channels I'm actually in.

Is this allowed?

This is the first question worth asking, so here's what I know. I'm not a lawyer, and none of this is legal advice.

In a lot of places the law is on your side about getting your own data. In the EU, GDPR Article 15 gives you the right to a copy of your personal data, and Article 20 says that, for data you provided, you get it in a "structured, commonly used and machine-readable format". The Digital Markets Act goes further for the biggest platforms and requires continuous, real-time portability. California's CCPA has a right to know, and the statute asks for the copy in a portable, readily usable format you can take somewhere else.

The law gives you the right to the data, though, not to any particular way of getting it. How you get it is the terms of service's business, and most of them say something about automated access. In the US, the Supreme Court's Van Buren decision made it a lot harder to argue that breaking terms of service is a computer crime when you're reading something you're allowed to see. Breaking them can still cost you the account, though.

My rules of thumb:

Ways to get your data out

Writing a connector is the last resort, not the first one. Roughly from easiest to hardest:

  1. The official export button. Google Takeout, ChatGPT's data export, "Download your information" on Facebook and Instagram. Usually a zip in your email within a few hours. Fine once. Hard to do every morning.
  2. Ask for it. A GDPR or CCPA access request by email. Slow, sometimes weeks, but it can include things no export button shows you.
  3. A service-to-service transfer. Some platforms move data straight to another one. The Data Transfer Initiative works on making this normal.
  4. The official API with your own token. GitHub's REST API is the easy case. It's stable and documented, and the token is something you create yourself.
  5. An archive tool. For Slack that's slackdump, which saves your channels, DMs and threads without admin rights. Slack's own export needs an admin, and on the Free and Pro plans it only covers public channels.
  6. Your signed-in browser. That's what this post is about. It's the most work to build, but once it works it's just "sign in once, collect every day".

The connectors repo has examples of most of these. The Claude connector requests the official export for you and reads the zip. GitHub uses the API with a token you create. The plain slack connector wraps slackdump. ChatGPT and the _browser ones read from a signed-in page, the way this one does.

Reading Slack the way Slack does

Open app.slack.com with the network tab open and you'll see the web client is mostly a loop of calls to /api/<method> on its own origin, carrying a token it keeps in local storage. Same Web API the docs describe, same method names.

So the connector opens its own browser window on the Slack sign-in page. I sign in with Google like I normally would. Once the session is live, the connector runs a small fetch inside that page, the same request the client would make. The token gets read inside the page and used there. It never comes out, and I never see it.

where the token stays signed-in app.slack.com tab local storage fetch /api/<method> token read and used here, never returned to the connector JSON only connector server
The connector only ever gets back the JSON Slack returns. The session stays in the browser where I signed in.

The first version worked on September 20. It passed the connector validator, and as a sanity check I ran the same query through Slack's own search and through the export and compared: 32 messages found by Slack, the same 32 in the export.

Two days later the old connector format got deleted from the connectors repo, which I'd known was coming but had hoped to ignore for a bit longer. So I rewrote it as a PDPP connector. The one decision I'm glad about: it reuses the record builders from the existing slackdump connector, so a message looks exactly the same whichever way it was collected. Anything reading the data can't tell, and shouldn't have to.

Two hours for a week

The rewrite was correct and painfully slow. It went through every conversation I'm in, about 420 of them, and paged back through each one until it hit the cutoff date. For one week of history that took somewhere between one and two hours, depending on the day.

It got worse once. I closed the browser window halfway through a run, the connector treated every failed request as something to retry, and it retried each remaining conversation six times against a page that no longer existed. Two hours, all of it wasted.

The fix for that one was easy: a closed page now ends the run right away. The slowness took more thinking. Most of those 420 conversations hadn't had a single message all week, and I was reading them anyway.

Asking Slack what changed

The web client already knows what changed, because it has to draw the unread badges. One call, client.counts, returns the time of the latest message in every open conversation. On the day I measured, 31 out of about 420 had anything newer than a week. Those are the only ones worth opening.

That leaves threads, where a reply on Thursday can hang off a parent from August. search.messages with an after: date returns everything written in the window, and each result's permalink carries the thread it belongs to. So the connector reads those threads directly, however old the parent is.

After that it was mostly tuning. History pages went from 200 messages to 999, and three conversations are read at once. I tried twelve parallel requests and didn't get rate limited, but search does push back about every ten pages, so I left it at three and stopped worrying.

one week of my Slack, wall-clock time read everything 1-2 h ask what changed 2m 40s 31 of ~420 conversations had a message that week
Same week, same data. The speedup is mostly not reading conversations nothing happened in.

What the data looked like

A few things surprised me once I had a week of my own Slack in a file.

Running it from the CLI

The connector isn't in the default catalog yet, so I needed vana-cli to run it from my own checkout, on schedule, like every other source. That's now vana connectors add slack_browser --from <checkout>, and after that vana collect --all picks it up.

Getting there turned up a couple of CLI bugs that had nothing to do with Slack. The CLI kills any connector that stays quiet for fifteen minutes, which is exactly what it does while you're signing in, so the connector now reports progress every minute while it waits. The CLI also rejected connector names with _browser in them, and it printed one line per message read, which for Slack meant thousands.

The shape of a connector

If you're thinking of writing one for a service you use, here's what this one turned into. Connectors live in PDP-Connect/data-connectors, one folder each, written in TypeScript. Slack ended up at about 2,300 lines of code and 1,600 of tests, which is more than I expected for "call some APIs and save the answers". Most of it is the boring parts: what to do when something fails, and where to resume tomorrow.

connectors/slack_browser/ lines manifest.json streams, options, needs a browser 1,171 index.ts entry: options, sign-in, run 341 web-api.ts the only code that touches the page 399 collector.ts what to read, cursors, concurrency 1,093 parsers.ts Slack JSON into records 231 schemas.ts validate every answer and record 223 types.ts options, conversation kinds 46 fixtures/api/ one recorded answer per API call 18 files fixtures/scrubbed/ real-shaped records, names removed *.test.ts run the whole thing against fixtures 1,632 outside the folder, by hand roster, orchestrator register the connector await-in-loop allowlist justify every await inside a loop
The two highlighted files are where the actual thinking went. Everything else is either declaration or plumbing.

The split that paid off is the one between web-api.ts and everything else. It's the only file that knows there's a browser. It runs the fetch inside the page and turns whatever comes back into one of a few outcomes: an answer, rate limited, session lost, page gone. The collector never sees Playwright. That's what makes it testable: the tests hand it a fake client that serves recorded JSON from fixtures/api/, named after the method that produced it.

how one message becomes a record slack tab web-api collector parsers schema check outcome which, how far shape or fail loudly
Each step only knows the one before it. Swap the tab for a fake and the rest runs unchanged in tests.

If you're building one

Some of this I knew going in. Most of it I learned by getting it wrong first.

  1. Start in the network tab. Before reading any API docs, watch what the site's own web app requests. If it calls a JSON API from its own origin, you can call the same thing from inside a signed-in page and never handle a password or token yourself.
  2. Find the change signal before you write the walk. Almost every app has some cheap way of knowing what's new, because it has to show you badges. Unread counts, an "updated since" filter, search by date. My first version walked everything and the fix was a rewrite of the planning, not a faster loop.
  3. Decide which failures stop the run. I ended up with three kinds. A closed page or a lost session stops everything right away. A rate limit means wait and try again. One conversation failing is a gap that gets recorded, and its cursor stays where it was, so tomorrow's run tries again. Retrying everything is how I got two hours of retries against a window I'd closed.
  4. Keep the cursor per thing, and read a little behind it. Mine stores the newest message read in each conversation, and each run starts a week below that. Replies and reactions land on old messages all the time, and a cursor that only moves forward misses them.
  5. Talk to the host. The CLI kills a connector that's been silent for fifteen minutes. Sign-in is silent. So is a long page walk. A progress line every so often costs nothing.
  6. Check yourself against the app. The most useful test I ran was the dumbest one: search for something in Slack, find the same thing in the export, compare the counts. It's how I found the join event, and how I knew the rest was right.
  7. Reuse what exists. If the source already has a connector, use its record builders even if you collect a completely different way. Then the data looks the same no matter who collected it, and nothing downstream has to care.
  8. Use the scaffolder, then do the wiring it skips. bin/connector-init.ts in packages/polyfill-connectors creates the folder. Registering it in the roster and orchestrator is manual, and the await-in-loop allowlist wants a file, line and column for every await inside a loop. Update those rows as the very last step, since any edit above them shifts the line numbers.
  9. Don't point the dev runner at your own Chrome. There's an environment variable that attaches the connector to an existing browser over CDP. It assumes the browser is a throwaway and closes every tab in it. I found this out with about forty tabs open.

The repo has a long authoring guide and a checklist next to it, and both are worth reading before the first line of code. To try your connector the way a user would, point the CLI at your checkout with vana connectors add my_source --from ~/code/data-connectors, then run vana collect my_source. From then on it runs on the schedule like any other source.

A prompt to start from

I built most of this with Claude Code, and the first prompt matters more than you'd think. Here's roughly what I'd start a new one with now. Change the first line and the rest mostly holds:

I want a PDPP connector for <service> in this repo (PDP-Connect/data-connectors).
It should collect my own data from my signed-in account, nothing I can't see in the app.

Before writing code:
1. Read packages/polyfill-connectors/docs/connector-authoring-guide.md and
   CONNECTOR-CHECKLIST.md, then look at connectors/slack_browser as a reference.
2. Find out how the <service> web app loads its data: which JSON endpoints it calls
   from its own origin, how it authenticates, and what it uses to know what changed
   (unread counts, "updated since", search by date). Tell me before going further.
3. Propose the streams, the cursor for each, and which failures should stop the run.

Then build it:
- Scaffold with bin/connector-init.ts, then wire the roster, orchestrator and the
  await-in-loop allowlist by hand (refresh the allowlist rows last).
- Keep all page access in one file; the collector gets a client it can fake in tests.
- Record real API answers as fixtures, scrubbed of names and ids.
- Pace requests like the web app does. A closed page or lost session ends the run.
- Report progress at least once a minute.
- Never set PDPP_*_REMOTE_CDP_URL against my own Chrome.

Done means: npm run verify passes, the conformance script passes, and a real run via
`vana connectors add <key> --from .` plus `vana collect <key>` matches what
the app itself shows me for a small sample I can check by hand.

The step I'd never skip is the second one. Making the model stop and report what it found in the network tab before it writes anything saved me from at least one wrong design.

What's still missing

It doesn't read canvases, per-channel stats, or read markers. Each would cost a call per conversation, and I haven't missed any of them. It also only goes back a week by default. I set it that way because the daily run is what I care about, and a full history is one environment variable away if I ever want it.

The pull request is still open, so for now you'd run it from a checkout the same way I do. Once it merges it should just show up as another source in vana connect.