← Culture Dive

Privacy Policy

What personal data we hold, why, and what you can ask us to do about it.

Culture Dive — Privacy Policy

DRAFT FOR LEGAL REVIEW. This is not legal advice. This document was written by reading the Culture Dive codebase, table by table and call by call. It describes what the system does today, including the places where it does nothing. No lawyer has reviewed it. Do not publish it until Spanish counsel has reviewed it and every [PLACEHOLDER: …] has been filled in with a real, checked fact.

Version: [PLACEHOLDER: version number and effective date — set at publication]


1. Who is responsible for your data

Our lead supervisory authority is the Spanish Agencia Española de Protección de Datos (AEPD).

2. Who this policy is about

Two very different groups. Read the one that applies to you.

This policy covers the Culture Dive product. [PLACEHOLDER: confirm whether it also covers the culturedive.ai marketing site, which has its own cookie and analytics arrangements that the product does not share. If it does, the cookie section below must be extended. Today culturedive.ai/privacy returns a 404.]


Part A — People who use Culture Dive

3. What we hold about you

Your account

WhatWhy
Email addressIt is your identifier and how we reach you
Full nameSo colleagues and our staff know who is who
Job titleContext for our staff; optional
An internal identity reference linking to our authentication providerTo connect your sign-in to your account
Two-factor status (whether you have enrolled, and when)Copied from the authentication provider so we can show it. It is a display copy; the authentication provider holds the authoritative record
A "sessions valid after" timestampSo "sign out everywhere" can take effect instantly

We do not hold your password, your authenticator secret, or your passkeys. Those are held by Supabase Auth (section 13). We never see them.

Your access

Invitations

If you were invited and never accepted, we still hold your email address. You are in this policy too. Section 12 explains what you can ask for.

Billing

Our record of what happened

We keep an audit log. Each entry records who did what, in which organisation and workspace, whether it succeeded, what it affected, a request reference and a short session reference.

Two honest details:

What our staff write

Reports are written by people. A report can mention a person — a client contact, an executive, a journalist. That text is stored, revisioned and, once published, immutable.

What we do not hold

4. Why we hold it, and our legal basis

[PLACEHOLDER: no lawful basis has been formally recorded anywhere for any processing. The column below is our reading of what each purpose would most plausibly rest on, not a decision anyone has taken. Counsel to confirm each row, and the result should be recorded in the record of processing activities, which also does not exist yet.]

PurposeLegal basis (proposed)
Creating and running your account so your organisation can use the serviceContract, Article 6(1)(b) — or our legitimate interest in providing the service the contract is with your employer, Article 6(1)(f)
Inviting people, and keeping a record of invitations sentLegitimate interest — running access control for a client's workspace, and being able to show who was given access and when
Security: authentication, two-factor, session revocation, audit loggingLegitimate interest — keeping a multi-tenant system safe; and legal obligation, Article 6(1)(c), for accountability records
Billing and reconciliationContract; and legal obligation for accounting records
Supporting you and responding to your requestsContract; legitimate interest
Improving the serviceLegitimate interest

Where we rely on legitimate interest, you can object. See section 12.

5. Cookies

The Culture Dive product sets three cookies. All of them are strictly necessary or directly requested by you, which is why there is no cookie banner. A banner here would be theatre.

CookieWhat it doesLifetime
sb-…-auth-token (and its chunks and code verifier)Your sign-in session. Without it you cannot stay signed inSession, refreshed by use
cd_csrfProtects the internal admin application against cross-site request forgeryEnds when you close the browser
cd_passkey_offerRemembers that you dismissed the "add a passkey" prompt, so we stop asking180 days

The first two are set with HttpOnly and Secure and are not readable by scripts, except cd_csrf, which must be readable by our own script for the protection to work. The third is only ever set when you click to dismiss the prompt, and only on a page you are already signed in to.

No third party can be contacted from a page in the portal: the content security policy allows connections and images only from our own origin, and a report is rendered with everything blocked.

[PLACEHOLDER: confirm the 180-day lifetime of the passkey-prompt cookie is acceptable. It is a user-interface preference, which is the category that is normally exempt from consent, but the exemption usually contemplates something shorter. Shortening it costs nothing.]


Part B — People who appear in content we collect

6. What this is, in plain terms

Culture Dive measures how a brand is talked about. To do that we collect public content and analyse it. That content can name you even though you have never used this product, never gave us anything, and have no relationship with us.

We think you are entitled to a straight account of that.

7. What we actually collect today

Today we collect one thing: the answers that AI answer engines give to our questions.

On a schedule, we ask ChatGPT, Perplexity and Gemini a fixed set of questions about a client's brand and category — questions like "what is this brand known for", "what criticisms come up about it", "who are its main competitors". We store the answer.

Those answers are free text written by a language model. They routinely name real people — founders, executives, journalists, critics, creators. When that happens, we are holding personal data about you that came out of a model, not out of anything you published.

Two things follow, and both matter:

What we do not collect today. Our system has entries for Instagram, TikTok, Reddit, YouTube, news RSS feeds and general web pages. No collector for any of them has been built. Nothing in this product currently fetches a social post, an article or a review. When the first one is built, this policy will be updated before it runs — including a proper legitimate interests assessment and a decision on how we give notice to the people whose content we collect.

[PLACEHOLDER: this section must be rewritten, not amended, when the first social or news collector ships. The change from "a model mentioned you" to "we collected your post" is a different processing operation with a different balance and a different notice duty.]

8. What we do with it

Items marked in our system as link-only are never sent to an AI model and never reproduced — only linked. Answer-engine output is not in that category, because it was generated in response to our own prompt.

9. Our legal basis, and the balance we have struck

We rely on legitimate interests (Article 6(1)(f)): the interest of a business in understanding how it is perceived in public discourse, and our interest in providing that service.

The balancing factors, as we see them:

Two honest qualifications on that last point.

[PLACEHOLDER: the 125-character cap is not enforced by the software. It exists as a comment in one migration file. Nothing truncates a quote today. It must be enforced in code before this policy is published, or the sentence must be deleted. A minimisation control that only exists in a document is not a minimisation control.]

[PLACEHOLDER: no legitimate interests assessment has been carried out or written down. The paragraph above is the argument; it is not the assessment. Counsel to produce a documented LIA before publication.]

10. Telling you we hold it

Article 14 of the GDPR says that when we get personal data from somewhere other than you, we should tell you. We are not able to do that for content collected at scale: we generally have no contact details for the people named in it, and finding them would mean collecting more data about them than we hold now.

We therefore rely on Article 14(5)(b) — that individual notice would take disproportionate effort — and we publish this policy instead. [PLACEHOLDER: counsel to confirm the Article 14(5)(b) position and whether anything more is expected, for example a machine-readable notice or a searchable index page. This has not been assessed.]

11. If you are named in something we collected

You can object. You do not have to explain why in any particular form, and you do not have to be a customer.

Write to [PLACEHOLDER: privacy contact address] and tell us:

What we will do: search for your name in the content we hold, tell you what we find, and remove it unless we have a compelling reason not to. We will answer within one month. If we need longer because the request is complex, we will tell you inside that month and explain why.

What we cannot do well yet, stated plainly:

[PLACEHOLDER: this is currently a manual process with no tooling behind it. There is no index over names in collected content, no search built for this purpose, and no supported way to remove a single item — the system keeps a per-client fingerprint ledger of every item ever seen, designed to outlive the content itself, so removing an item without also handling that ledger would let the same item be re-collected as new. Before this section is published, build: a name search over collected content, a single-item removal path, and a tombstone in the fingerprint ledger so a removed item stays removed. Until then, the commitment above can be honoured, but slowly and by hand, and the policy should say so rather than imply a self-service process exists.]

You can also complain to the AEPD, or to the supervisory authority where you live. Details are in section 15.


Both groups

12. Your rights, and what we can actually do today

You have the right to ask for a copy of your data, to have it corrected, to have it deleted, to have its use restricted, to object to processing based on legitimate interests, and to receive it in a portable format. You will never be charged or penalised for asking.

What you can do right now, yourself, in the portal:

RightAvailable today
See what we hold about your account — email, name, job title, two-factor status, your workspaces and rolesYes, on your profile page
Correct your name and job titleYes, on your profile page. You can change only these two fields; the design deliberately prevents anyone changing anything security-relevant about their own account
End every session everywhereYes
Add or remove two-factor authentication and passkeysYes

Everything else is a manual request to [PLACEHOLDER: privacy contact address]. We will answer within one month.

Where we have to be honest about limits:

Whatever we do in response to a request, we record it in the audit log.

13. Who else handles your data

We keep this list short on purpose, and it lists only providers that are actually wired into the running system.

ProviderWhat it handlesWhere it is processedCompany base
SupabaseThe database — everything in Part A and Part B — plus authentication, two-factor factors and passkeysEU, region eu-central-1 (Frankfurt)United States
Fly.ioRuns the application and the background workers; application logsFrankfurt (fra), all four applicationsUnited States
PostmarkSends invitation and notification email; receives bounce notificationsUnited StatesUnited States
StripePayments. Holds the payer's name, email and billing address; sends us event payloads[PLACEHOLDER: confirm which Stripe entity contracts with APC Labs and where the account's data is processed]United States / Ireland
OpenRouterThe gateway through which every AI call is made, and the models behind it[PLACEHOLDER: unknown — see below]United States
SentryError reports from the applicationEU region, enforced in code — the application refuses to start against a non-EU Sentry hostUnited States

About OpenRouter and the models behind it. OpenRouter routes each request onward to a model provider. Our system deliberately does not pin which provider serves a call — that portability was a design decision — with the consequence that we cannot say from our own records which company processed any given request. [PLACEHOLDER: this is a real problem for the record of processing activities and for the sub-processor list a client will ask for. Two things must happen: OpenRouter's account-level data policy (retention, logging, and whether prompts may be used for training) must be checked, tightened and written down — nothing in the code sets any of it, so whatever the account default is, is our position; and a decision is needed on whether to record the serving provider per call.]

Not on this list, and worth saying: we use no third-party analytics, no advertising network, no marketing automation platform, no customer-data platform, and no data-broker or scraping vendor. Some such vendors appear in our internal planning documents. None is connected to anything.

14. Sending data outside the EU

All processing we control happens in the EU: the database is in Frankfurt, the application runs in Frankfurt, error reporting is on an EU-region host and the code refuses to start otherwise.

But several of our providers are US companies, and one of them — the AI gateway — routes requests onward to providers whose location we do not record. So we cannot claim there is no international transfer.

[PLACEHOLDER: for each provider in section 13, counsel must establish the transfer basis and record it — adequacy decision (for example the EU–US Data Privacy Framework, and whether the provider is certified), standard contractual clauses, and whether a transfer impact assessment is required. This has not been done for any of them. The OpenRouter position is the hardest and should be settled first, because it is the one where the receiving party is not fully known.]

15. How long we keep things

We have to be straightforward here. This system does not currently delete anything, and no retention schedule has been set.

The database is built for it — the two largest tables are partitioned by month specifically so that old data can be dropped — but the job that would drop them has never been written. Every category below therefore has, today, an actual retention period of "indefinitely".

DataWhat we intendWhat happens today
Account recordsKept while your organisation is a client, then blocked for the statutory liability period, then deleted [PLACEHOLDER: period]Kept indefinitely. Nothing deletes and nothing blocks
Invitations, including those never accepted[PLACEHOLDER: period — this one is easy and should be short. An unaccepted invitation past its 48-hour expiry has no remaining purpose]Kept indefinitely
Audit log[PLACEHOLDER: period. It is the accountability record, so it should outlive most other categories, but not forever]Kept indefinitely
Collected content and the AI readings of it[PLACEHOLDER: period. This is the most urgent one: it is third-party personal data held under legitimate interests, where an unbounded period is the hardest position to defend]Kept indefinitely
Published reportsKept as a permanent record of what was delivered [PLACEHOLDER: confirm — reports are immutable by design, so a name that reaches one stays in the record]Kept indefinitely
Billing records and stored payment-provider payloadsThe statutory accounting period [PLACEHOLDER: period under Spanish law]Kept indefinitely
Backups[PLACEHOLDER: the managed database provider keeps daily backups for 7 days. Point-in-time recovery is not purchased. A separate backup script writes database dumps to a location set by an environment variable, with no stated retention, no stated encryption and no stated geography. All three need deciding and writing down]See left

[PLACEHOLDER: a retention schedule per data class is the single most important gap in this document, and it blocks the document. A privacy policy that cannot state a retention period is not finished. Two engineering traps for whoever builds the job: the per-client content fingerprint ledger is deliberately not partitioned and will survive any partition drop, so it needs its own pruning schedule; and dropping a partition is a schema operation that will orphan the AI readings attached to it rather than removing them.]

16. AI and automated decisions

Nothing we send to an AI model contains your account data. We checked this specifically. When we ask an answer engine a question, the request contains our own prompt, the brand name and the category, and nothing else — no user record, no email address, no account identifier, no IP address. The optional field that would let us tag a request with a user identifier is available to us and is deliberately never used.

Content we have collected does go to a model. That is the reading step described in section 8. If a collected item names a person, that name goes to the gateway and onward to a model provider.

No decision about you is made automatically. There is no scoring, ranking or profiling of individuals in this product. Sentiment attaches to a piece of content, not to a person — there is no person record for anyone in Part B and nothing is aggregated against a name. Nothing in this product produces a legal effect or a similarly significant effect on any individual, so the automated-decision rules in Article 22 do not apply.

A person rules on every conclusion. AI output is a candidate. A member of our staff keeps it, kills it or merges it, and a named person authors and publishes every report.

If that changes, this section changes first. Specifically: if we ever build a feature that identifies, ranks or evaluates named individuals — for example mapping creators by their affinity to a brand — that is profiling, it needs its own assessment under both the GDPR and the EU AI Act, and it will not ship before that assessment is written. [PLACEHOLDER: this feature appears in the commercial design materials as a priced product. It does not exist in the software. If it is going to be built, the assessment should be commissioned now rather than during the build.]

17. How we protect data

18. Changes to this policy

We will update this policy when the product changes, and we will change it before the change goes live, not after. Where a change materially affects you, we will tell you.

[PLACEHOLDER: there is currently no mechanism to show a user a legal document or to record that they have seen one. Two database tables were specified for this and neither exists, and there is no screen in the product on which a document could be shown or accepted. Until it is built, "we will tell you" means by email, and this paragraph should say so.]

19. Complaints

If you are unhappy with how we have handled your data, tell us first at [PLACEHOLDER: privacy contact address]. We would rather fix it.

You can also complain to the Spanish data protection authority:

> Agencia Española de Protección de Datos (AEPD)

> C/ Jorge Juan, 6, 28001 Madrid, Spain

> www.aepd.es

If you live in another EU country, you can complain to your own supervisory authority instead.


Reviewer notes — Privacy Policy

Everything the draft above cannot yet state as fact, with the evidence.

Blocking. This policy cannot honestly be published until these are resolved.

#ItemEvidence
P1No retention schedule exists and nothing deletes anything. The worker runs three scheduled tasks — sweep, heartbeat, reconcile — and none is a retention job. Signal and audit tables are partitioned monthly *for* retention; no code drops a partition. The repository already records this twice without acting on it.apps/api/app/workers/tasks.py:173,263,310; 0023_partitions.sql:54-56,158-160; docs/00-ASSESSMENT.md:138; docs/04-OPEN-QUESTIONS.md:156
P2A user cannot be deleted, blocked or exported. No route (/me has four endpoints, none deletes); no runtime role holds DELETE on app.app_user; twenty foreign keys reference it with zero CASCADE or SET NULL, two of them explicit RESTRICT; deleted_at/blocked_at are read on the auth path and written by nothing; data_export and data_deletion exist in the audit vocabulary with nothing able to emit them.apps/api/app/main.py:69,132,148,175; 0002_tenancy.sql:54-55,64,77,220; 0012_self_service_columns.sql:47; apps/api/app/core/authz.py:82-83,126; 0001_foundation.sql:119
P3The 125-character cap is not implemented. One comment on an enum value. signal.body is unbounded text. Nothing truncates. It is cited as the minimisation control in three internal documents.0016_signals.sql:41,128; CONVENTIONS.md:75; docs/06-DESIGN-SYSTEM.md:160; docs/04-OPEN-QUESTIONS.md:35
P4No lawful basis has ever been recorded, and no LIA exists. Section 4 and section 9 are proposals, not decisions. There is no ROPA.Nothing in the tree records one; docs/09-BUILD-BACKLOG.md:243 lists ROPA/DPA/DPIA/retention schedule as needing counsel
P5No transfer assessment for any processor. OpenRouter is the hard case: models, provider and route are banned as vendor fields by architectural decision, so the serving provider cannot be pinned or recorded, and no data-policy or logging parameter is set — the account default is the company's position by omission.apps/api/app/integrations/llm.py:46-50,62-76; apps/api/app/integrations/openrouter.py; docs/04-OPEN-QUESTIONS.md:152
P6The DSAR-against-collected-content path in section 11 has no tooling. No name index, no search, no single-item removal. app.signal_dedupe is unpartitioned by design and holds a per-tenant SHA-256 fingerprint of every item ever seen; deleting a ledger row lets the same item be re-collected as new, so removal needs a tombstone, not a delete.0016_signals.sql:167-180; apps/api/app/services/signals.py:35-48

Corrections this review made to the draft, which the reviewer should not undo

#ItemEvidence
P7There is no social or news collector. Nine sources are registered — Instagram, TikTok, Reddit, YouTube, news RSS, web — and nothing fetches from any of them. The only writer of app.signal anywhere in the tree is the answer-engine probe, where author_handle holds the engine key, not a person. The Part B exposure today is model output that names people, which is an accuracy problem as much as a transparency one. Any internal claim that third-party scraping is running "right now, at scale" is wrong.0016_signals.sql:249-261; apps/api/app/services/aeo.py:225,288-307; apps/api/app/integrations/ contains llm, openrouter, postmark, stripe, supabase_admin and nothing else
P8The audit log's ip and user_agent columns are never written. The only INSERT lists thirteen columns and neither is among them. The policy says so rather than letting the schema imply otherwise.0003_audit.sql:37-38; apps/api/app/services/audit.py:69-77
P9The audit append-only trigger no longer exists. It was created on the default partition in 0003, and 0023 drops and recreates that partition inside a loop. No migration restores it; the eighteen monthly partitions never had it; no test attempts an UPDATE or DELETE. Immutability now rests solely on withheld grants. Separately, seq, prev_hash and hash are NULL on every row and app.audit_chain_head has no writer — append-only, not tamper-evident.0003_audit.sql:99-117,73-80; 0023_partitions.sql:141-144; apps/api/app/services/audit.py:1-9; apps/api/tests/sql/isolation_proof.sql:309-340
P10Email tracking is genuinely off and tested. TrackOpens: False, TrackLinks: "None", both asserted in CI. No analytics vendor appears in either frontend's dependencies. Sentry is EU-enforced with PII off and an aggressive scrubber, and is required to boot in production. This is the strongest control in the tree and the policy states it positively.apps/api/app/integrations/postmark.py:141-142; apps/api/tests/integrations/test_postmark.py:87,98; apps/api/app/core/observability.py:35,57-101,114-124,135; apps/api/app/core/config.py:138
P11No user personal data reaches any LLM. The probe request is a rendered prompt plus brand and category; the OpenAI user / safety_identifier fields are permitted by the client and never passed. Worth stating positively — most policies cannot.apps/api/app/services/aeo.py:49-52,108-115,249-253; apps/api/app/integrations/llm.py:31,33,169-174
P12Three cookies, and no banner is the correct answer. Verified by finding every cookies() call site in both applications. The API sets none. The portal's content security policy is img-src 'self' data: / connect-src 'self', and a rendered report is default-src 'none'. Stripe is a hosted redirect with no client-side Stripe dependency.apps/portal/lib/supabase/server.ts:50; apps/portal/proxy.ts:54; apps/admin/app/(admin)/csrf-provider.tsx:24; apps/portal/app/(portal)/passkey-offer-actions.ts:26-35; apps/portal/lib/headers.ts:53,55; apps/portal/app/api/render/route.ts:56
P13culturedive.ai/privacy returns 404 and /terms §5 states the opposite of this document — that no personal data is collected by scraping and that the public data analysed is not personal data. That is a legal conclusion published on the company's own site which app.signal.author_handle contradicts on its face. It also names two different controllers in one document. Correct the site in the same change that publishes this policy.Live site, checked; 0016_signals.sql:113-135
P14The stored Stripe payloads are named explicitly in the policy because the migration that created the table names the risk itself: *"it holds Stripe payloads with customer emails and amounts across every tenant."* The webhook stores the raw wire body.0009_commerce.sql:169,192; apps/api/app/hooks.py:199-203
P15Section 12's self-service rectification claim is accurate and unusually well built. PATCH /me takes no user id at all, and the database grant is column-scoped to full_name and job_title, closing an escalation where a client could previously clear their own blocked_at or repoint their identity binding.apps/api/app/main.py:148,155-157; apps/api/app/services/profile.py:111; 0012_self_service_columns.sql:4-25,51