KOLCO CLUB · Training data for design models

Design data, drawn by people.Licensed for machines.

3,096,618 editable SVG design files, shipping today. Every one carries live text, its own metadata row and a documented origin — and no generative model made any part of any of them. Licensed to AI companies for training, and open to creators who want their work in it.

Live editable text Per-file metadata No AI-generated artwork Self-contained SVG
3,096,618editable SVGs
shipping today
996,618poster & print
templates
1,800,000infographic
templates
300,000residential
floor plans
the doodle desk — marketplace sample — 191 files 4 of 191 shown
§ 01What it trains

Four things design models are bad at.

Nobody wakes up wanting poster data. They wake up because a model cannot render legible text, or cannot produce an editable vector, or cannot lay out a page. These are the same three million files, cut along the axis of what they teach.

Vectorisation

Image in, real SVG out

Every asset is a vector source that can be rendered deterministically to raster at any size — so every file is an aligned (image, ground-truth SVG) pair with clean geometry, not a trace.

Built from
the whole corpus
Pairs
~3,096,618
Status
Derivable on request
Text in images

Legible type, with ground truth

Text rendering is a known weak spot, and scraped data has no ground truth for it. Here every string is a live SVG <text> node with its exact content, position, size and colour — readable before a pixel is drawn.

Sample fact
10,866 in 191 files
Corpus est.
~169,000,000
Status
Derivable on request
Layout & composition

Where things go, and why

Every text node carries an explicit font size and position, and every page uses more than one size — so typographic hierarchy and reading order are recoverable per file, across 23 print formats and 587 layout families.

Built from
996,618 print pages
Verified
92/92 in sample
Status
Derivable on request
Charts & documents

Figures that agree with themselves

On any infographic the share breakdown, ranking, trend and comparison are four views of one set of numbers. A page is internally consistent before you change it, and still consistent after — which is what makes it usable for chart reasoning rather than chart mimicry.

Built from
1,800,000 infographics
Coverage
56 sectors, 896 subjects
Status
Shipping

“Derivable” means exactly that. The pairs, the text boxes and the hierarchy are all present in the files and extractable by a deterministic pass — we have verified that on the public sample and you can too. They are not pre-packaged SKUs sitting on a shelf. Tell us which one you need and we will cut it against your schema.

§ 02Catalogue

Two design collections, already drawn.

Shipping collections from The Doodle Desk, not a roadmap. Every file is real SVG with live editable text — no outlined type, no flattened artwork — and every asset carries its own metadata row, so a lab can license a slice instead of running a scrape.

Posters
& print

996,618editable SVGs
Story countdown poster
App card flyer
Big-type poster
Edge frame poster

Single-page print products at known trim sizes — the headline, body copy, call to action and contact line all live text.

  • Illustration posters & print — 498,659
  • Brand posters — 286,613
  • Poster & full-page ad — 206,313
  • Festival posters — 5,033
  • 205 business niches
  • 23 print formats
  • 56 typographic themes
  • Dublin Core RDF on every file
SVGLive textLibre fonts

Info­graphics

1,800,000editable SVGs
Fleet infographic
Membership infographic
Clinic infographic
Warehouse infographic

Charts, rankings and breakdowns where the arithmetic on a page agrees with itself — before and after you change a figure.

  • Composite — 700,000
  • Slide-deck, 16:9 — 500,000
  • Poster-size, A2 — 350,000
  • Indian business — 250,000
  • 896 subjects across 56 sectors
  • 16 page formats
  • 1000×1000 to 1100×3000
  • Chart forms recorded per file
SVGLive textA2 / A3 / 16:9
§ 03Proof

Measured, not asserted.

Every figure below was run against the 191 SVG files in the public review sample — stratified, drawn at random within each format, nothing retouched. Download it and reproduce the table yourself.

191 sample files · measured on the files themselves
CheckPostersInfographicsFloor plans
Files in sample924851
Parse as valid XML92 / 9248 / 4851 / 51
Carry live, non-empty <text>92 / 9248 / 4851 / 51
Reference a remote font or asset000
Median file size88 KB76 KB41 KB
Median live text nodes761151
Median vector paths2032275

Fonts travel inside the file as data URIs, or are named system faces with generic fallbacks. Nothing reaches outside itself for a font or an image, so a file renders identically wherever it lands — including inside a training pipeline with no network. The full breakdown — class distributions, the label schema, worked examples and the provenance record — is in the datasheet.

§ 04Sample

Take the sample before you talk to us.

191 files with their metadata indexes, catalogue indexes and licence notes. No form, no email, no call. Open them in any vector editor and judge them.

These files are for evaluation, not redistribution. The commercial terms travel with the full pack — ask and we will send the licence before any conversation about price.

§ 05Licensing

Two ways in.

Evaluate for free, then licence what you need. Scoping goes down to the sub-collection — icon sets alone, or floor plans alone — so you are never obliged to take the whole catalogue to get the part you need.

Free, no contract

Evaluate

The public sample plus, on request, a larger evaluation slice cut to your taxonomy. For assessment only — no training grant.

  • 191 files, downloadable now
  • Full record schema and manifests
  • Larger slices on request
Annual, non-exclusive

Licence a release

A versioned release of one or more collections, licensed for model training, with the documentation your model card needs. Need something the catalogue does not hold? That gets drawn to your brief under the same terms.

  • Scoped by collection or sub-collection
  • Quarterly refresh with new assets
  • Consent manifest and QA report
  • Commissioned sets to your taxonomy
Delivery

Your S3 or GCS bucket, a private Hugging Face repo, or signed archive links.

Packaging

WebDataset shards or Parquet, sources and renders addressed by checksum.

Versioning

Immutable release tags. Additions ship as new tags; withdrawals are listed.

Paperwork

Consent manifest, QA report and a provenance summary written for a model card.

§ 06Creators

Submissions are open.

The catalogue today is our own production. We are opening it to other designers: submit original vector work, we check and label it, and it joins the releases that AI companies licence. You are paid a share of what each release earns, for as long as it earns.

What stays yours

  • Copyright. You own the work before and after. KOLCO only ever licenses it.
  • Every other channel. Sell the same files on your store, on stock sites, to clients.
  • The right to withdraw. Pull files from all future releases at any time.
  • Category control. Allow illustration but not brand marks. Allow icons but not faces.
  • Attribution on request. Be named in the release credits, or stay anonymous.
  • A clear ledger. Which release, which files, which buyer, how much.
§ 07Questions

The questions both sides ask first.

Do creators give up their copyright?

No. You keep full ownership. What you grant is a non-exclusive licence for the work to be included in dataset releases used for model training. You can keep selling, printing and licensing the same files anywhere else, at the same time.

How are creators actually paid?

Each release earns licence revenue. A fixed share of that revenue is distributed across the creators whose files are in the release, weighted by how many of their accepted files it contains. Payouts run monthly, and every line is traceable to a release and a buyer.

Can I remove my work later?

Yes, for anything that has not shipped. Withdrawing removes your files from all future releases and refreshes. Releases already licensed remain valid for their term, because a buyer has already trained on them — that limit is stated plainly in the grant you sign, before you submit.

Can a lab license just one category?

Yes. Scoping goes down to the sub-collection — icon sets alone, or floor plans alone, or festival illustration alone. You are not obliged to take the whole catalogue to get the part you need.

Can we inspect the data before licensing anything?

Yes, and you should. A 191-file review sample is published for the three shipping collections — stratified rather than cherry-picked, drawn at random within each format, with nothing retouched. It ships with the per-file metadata index, the catalogue index and the licence note, so you can measure the claims on this page yourself before a contract exists.

Was any of this generated by an AI model?

No. No generative model produced any part of any file in the shipping packs, and the licence note that travels with each pack states it. That matters more than it used to: a design model trained on another model’s output inherits its artefacts, and there is no way to detect that after the fact from the files alone.

Get in touch

Tell us which side you are on.

AI companies

Licence a release

Tell us what you are training and where your current data falls short. We reply with the record schema and an evaluation slice from the closest collections.

Creators

Submit your work

Send a portfolio link and roughly what is in your archive. We come back with the categories that fit and what a submission looks like.