AI training & evaluation data

Real human interaction data for models that need social context

Custom datasets built from consent-based dating-app exports, with profile context, longitudinal behavior, conversation structure, and approved redacted text.

Model-training and evaluation rights are scoped through a custom agreement. Field availability depends on the approved use case.

12,000+
profiles
294M
swipes
3.1M
matches
1.1M
messages
custom-ai-dataset.jsonl
approved fields · sample record
{
"profile": {
"ageAtUpload": 27,
"country": "NO"
},
"conversation": {
"primaryLanguage": "en",
"totalMessageCount": 18,
"responseTimeMedianSeconds": 420
}
}
Source
GDPR exports
Format
JSONL
PII
Redacted
License
Custom
Built for model teams

A domain dataset with real context

SwipeStats supports custom training and evaluation work where conversational structure, provenance, and behavioral context matter.

Post-training

Adapt models to real social conversation

Use naturally occurring dialogue structure, message direction, timing, and language context to develop systems that handle human conversation with more nuance.

Evaluation

Build domain-specific tests

Create held-out evaluations for social reasoning, conversation continuity, ambiguous intent, and privacy-preserving behavior in a high-context domain.

Trust & safety

Study sensitive interaction patterns

Develop approved datasets for moderation, redaction, and safety research with explicit provenance and a documented handling boundary.

Behavior modeling

Connect language to longitudinal context

Pair conversation structure with activity, match history, response timing, and profile context for richer modeling and analysis.

Dataset layers

Scope the fields around the model objective

Every delivery begins with a field review. The resulting schema includes the minimum data required for the approved training or evaluation workflow.

Inspect the public sample
01

Profile context

Age, geography, interests, preferences, and account history

02

Longitudinal behavior

Daily usage, swipes, matches, activity windows, and response timing

03

Conversation structure

Message order, direction, language, duration, and engagement metadata

04

Approved text fields

Redacted profile and message text where the engagement allows it

Provenance & governance

A documented data boundary for sensitive human context

Dating conversations need a higher handling standard. We scope each engagement around provenance, privacy controls, intended model use, and delivery requirements.

Source data comes from user-submitted official dating-app exports
Pseudonymous identifiers replace platform and account identifiers
PII-redacted fields are used for standard text deliveries
Allowed model uses, retention, and redistribution are defined in writing
Dataset limitations and field coverage accompany the delivery
Final suitability, legal terms, and security requirements are reviewed for each customer and intended model use.
Custom delivery

From model objective to usable dataset

01

Define the model objective

Share the intended training, evaluation, safety, or research workflow.

02

Review fields and coverage

We confirm available cohorts, languages, time ranges, text fields, and sample size.

03

Agree the data boundary

The license records permitted model uses, retention, security, and redistribution terms.

04

Receive a documented export

Delivery includes versioned JSONL, a schema guide, coverage notes, and citation metadata.

Questions

Before you scope a dataset

Can we use the data to train or fine-tune a model?

Potentially. AI training and evaluation are licensed through a custom agreement that names the approved model use, fields, retention period, and redistribution boundary.

Does the dataset include conversation text?

Conversation structure and derived metrics are broadly available. Redacted message text can be considered for approved custom engagements after a field and privacy review.

Does a delivery contain raw personal information?

Standard text deliveries use redacted fields and pseudonymous identifiers. The final schema and handling requirements are documented before delivery.

Can we request a particular cohort or language?

Yes, subject to coverage and privacy thresholds. We inventory the requested slice before confirming a dataset size or delivery date.

Custom AI datasets

Bring us the model objective

We’ll confirm the available fields, privacy boundary, sample size, and licensing path for your training or evaluation work.