| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
Torque is a declarative, fully typesafe DSL for quickly building complex LLM synthetic datasets. Compose conversations like components, generate realistic variations with any model efficiently.
import * as T from "@qforge/torque";
import { openai } from "@ai-sdk/openai";
await T.generateDataset(
() => [
T.generatedUser({ prompt: "Friendly greeting or introduction" }), // AI generated
T.oneOf([
// pick one randomly (weights are optional)
{ value: T.assistant({ content: "Hello!" }), weight: 0.3 }, // static
T.generatedAssistant({
prompt: "Respond to greeting",
reasoning: T.generatedReasoning({
prompt: "Reason about the greeting",
}),
// or reasoning: reasoning({ content: "...." }),
}), // AI generated, gets remaining weight
]),
T.times(between(1, 3), [
T.generatedUser({
prompt: "Chat about weather. Optionally mentioning previous message",
}),
T.generatedAssistant({ prompt: "Respond to user. Short and concise." }),
]),
],
{
count: 2, // number of examples
model: openai("gpt-5-mini"), // any ai-sdk model
seed: 42, // replayable RNG
metadata: { example: "quick-start" }, // optional per-row metadata
}
);Outputs:
{"messages":[{"role":"user","content":[{"type":"text","text":"Hi there! I'm new here and just wanted to say hello."}]},{"role":"assistant","content":[{"type":"text","text":"Hello!"}]},{"role":"user","content":[{"type":"text","text":"The sunshine today is perfect for a walk in the park."}]},{"role":"assistant","content":[{"type":"text","text":"Absolutely—warm and bright out there."}]},{"role":"user","content":[{"type":"text","text":"Do you think the clouds will roll in later this evening?"}]},{"role":"assistant","content":[{"type":"text","text":"Maybe briefly, but it should stay mostly clear."}]}]}
{"messages":[{"role":"user","content":[{"type":"text","text":"Hey! Hope you're having a great day."}]},{"role":"assistant","content":[{"type":"text","text":"Hi there! I'm doing great—what can I help you with?"}]},{"role":"user","content":[{"type":"text","text":"The weather keeps flipping between sun and drizzle lately."}]},{"role":"assistant","content":[{"type":"text","text":"Totally—it’s been bouncing around all week."}]},{"role":"user","content":[{"type":"text","text":"Should I expect rain again tonight?"}]},{"role":"assistant","content":[{"type":"text","text":"Pack an umbrella just in case; there’s a chance of showers."}]},{"role":"user","content":[{"type":"text","text":"Thanks! I’ll be prepared if it turns stormy."}]},{"role":"assistant","content":[{"type":"text","text":"Good call—better to stay dry than sorry."}]}]}💡 See full example: examples/quick-start.ts | ▶️ Try in Browser
Building synthetic datasets for LLMs is tedious:
Torque solves this with a declarative approach. Just like React transformed UI development from imperative DOM manipulation to composable components, Torque transforms dataset generation from manual JSON editing or writing complicated scripts to declarative conversation schemas. Plus, its optimized structure means you can use smaller, cheaper models while benefiting from cache optimization for lower costs.
npm install @qforge/torque
# or
bun add @qforge/torqueBuild conversations by composing message schemas, you can compose them together to build complex conversations from reusable parts:
// Reusable greeting pattern
const greeting = () => [
system({ content: "You are a helpful assistant." }),
user({ content: "Hello!" }),
assistant({ content: "Hi! How can I help?" }),
];
// Compose it with additional conversation
const extendedSchema = () => [
...greeting(),
user({ content: "What's the weather like?" }),
assistant({ content: "I'd be happy to check that for you!" }),
];
// Or create variations
const formalGreeting = () => [
system({ content: "You are a professional assistant." }),
user({ content: "Good morning." }),
assistant({ content: "Good morning. How may I assist you today?" }),
];
const schema = () => [
// Weighted selection between schema branches
oneOf([
{ value: greeting(), weight: 0.6 },
formalGreeting(),
extendedSchema(),
]),
// Continue with shared conversation flow
generatedUser({ prompt: "Ask a question" }),
generatedAssistant({ prompt: "Provide helpful answer" }),
];💡 See full example: examples/schema-composition.ts | ▶️ Try in Browser
Use metadata({ ... }) inside your schema to hoist custom fields into the generated row. The helper runs during the check phase, so you can safely compute metadata once and reuse it in generation. Passing a function receives the current metadata object (merged with any top-level metadata you provided) and may mutate it or return a new object for advanced scenarios like simple counters.
const schema = () => [
system({ content: "You are a helpful assistant." }),
oneOf([
() => [metadata({ variant: "static" }), assistant({ content: "Hello!" })],
() => [
metadata({ variant: "generated" }),
generatedAssistant({ prompt: "Greet the user warmly" }),
],
]),
];
const withCounter = () => [
metadata((meta) => {
meta.count = meta.count ?? 0;
meta.count += 1;
count += 1;
}),
generatedUser({ prompt: "Ask a question" }),
];When the dataset is saved, you can read these values under row.meta.metadata.
When generating datasets with a seed, each row automatically receives a unique, deterministic id in its metadata. This ID is generated based on the seed value, making it easy to identify and track specific rows across multiple runs.
await generateDataset(schema, {
count: 3,
seed: 100,
model: openai("gpt-4o-mini"),
});
// Output in row.meta.metadata:
// Row 0: { id: "row_100_1l2dpno" } // seed: 100
// Row 1: { id: "row_101_txnff9" } // seed: 101
// Row 2: { id: "row_102_2sx56u" } // seed: 102The ID combines the seed value with a deterministic hash, ensuring:
If custom metadata is provided, the ID is automatically merged with it:
await generateDataset(schema, {
count: 2,
seed: 100,
model: openai("gpt-4o-mini"),
metadata: { projectName: "my-project" },
});
// Output: { id: "row_100_1l2dpno", projectName: "my-project" }Note: IDs are only generated when a seed is provided. Without a seed, no ID is added.
Build dynamic, varied datasets with composition helpers:
import { oneOf, times, between, optional } from "@qforge/torque";
const schema = () => [
// Choose randomly from options (weights optional)
oneOf([
user({ content: "Hello" }),
{ weight: 0.5, value: user({ content: "Hi there" }) },
user({ content: "Hey" }),
]),
// Repeat pattern 3 times
times(3, [
generatedUser({ prompt: "Ask a question" }),
generatedAssistant({ prompt: "Answer the question" }),
]),
// Repeat random number of times (1-5)
times(between(1, 5), [generatedUser({ prompt: "Follow-up question" })]),
// Optionally include (50% chance)
optional(assistant({ content: "Anything else I can help with?" })),
];oneOf accepts plain schema entries or { value, weight } objects. Provide any subset of weights (summing to ≤ 1) and the remaining probability is spread evenly across unweighted entries.
Pass a uniqueBy configuration when you need each option to be used at most once across every row/schema during generation. When using uniqueBy, each option must be an object with id, value, and optionally weight:
const toolOptions = [
{ id: "weather", value: weatherTool.toolFunction() },
{ id: "calendar", value: calendarTool.toolFunction() },
{ id: "flight", value: flightTool.toolFunction() },
];
const schema = () => [
oneOf(toolOptions, {
uniqueBy: {
collection: "tools",
},
}),
];You can also combine uniqueBy with weighted options:
const toolOptions = [
{ id: "weather", value: weatherTool.toolFunction(), weight: 0.5 },
{ id: "calendar", value: calendarTool.toolFunction(), weight: 0.3 },
{ id: "flight", value: flightTool.toolFunction(), weight: 0.2 },
];
const schema = () => [
oneOf(toolOptions, {
uniqueBy: {
collection: "tools",
},
}),
];The collection name identifies the shared pool (so multiple oneOf calls can coordinate). The id property must be a string, number, or boolean and is used to track uniqueness. Torque throws if the pool is exhausted, making it easy to guarantee perfect round-robin coverage.
For a simpler API, use uniqueOneOf to automatically generate IDs and create a reusable function. This is especially useful when you want to create the unique selection function outside of your schema:
import { uniqueOneOf } from "@qforge/torque";
// Create the factory function outside generation
const tools = [weatherTool, calendarTool, flightTool];
const oneOfTools = uniqueOneOf(tools);
// Or with weighted options
const weightedTools = [
{ value: weatherTool, weight: 0.5 },
{ value: calendarTool, weight: 0.3 },
flightTool, // unweighted, gets remaining weight
];
const oneOfWeightedTools = uniqueOneOf(weightedTools);
const schema = () => {
const tool = oneOfTools(); // Returns a unique tool each time
return [
tool.toolFunction(),
generatedUser({ prompt: "Ask question requiring this tool" }),
generatedToolCall(tool, "t1"),
generatedToolCallResult(tool, "t1"),
];
};The uniqueOneOf factory automatically:
💡 See weighted example: examples/weighted-one-of.ts
💡 Full utilities demo: examples/composition-utilities.ts | ▶️ Try in Browser
Define tools with Zod schemas for complete type safety:
import {
tool,
generatedToolCall,
generatedToolCallResult,
} from "@qforge/torque";
import { z } from "zod";
// use standard tool schema using zod ensuring complete type safety
const weatherTool = tool({
name: "get_weather",
description: "Get current weather for a location",
parameters: z.object({
location: z.string().describe("City name"),
units: z.enum(["C", "F"]).optional(),
}),
output: z.object({
temperature: z.number(),
condition: z.string(),
}),
});
const schema = () => [
weatherTool.toolFunction(),
generatedUser({ prompt: "Ask about weather in a city" }),
generatedToolCall(weatherTool, "t1"), // type safe 100% correct generated tool calls
generatedToolCallResult(weatherTool, "t1"), // similarly 100% correct generated tool results
generatedAssistant({ prompt: "Interpret the weather data for the user" }),
];💡 See full example: examples/tool-calling.ts | ▶️ Try in Browser
Torque is built with TypeScript and provides complete type safety. Both for user and AI generating the data. Ensure that the arguments and tool results are always matching schema.
// Full type inference for tool parameters
const weatherTool = tool({
name: "get_weather",
description: "Get current weather for a location",
parameters: z.object({
location: z.string().describe("City name"),
units: z.enum(["C", "F"]).optional(),
}),
output: z.object({
temperature: z.number(),
condition: z.string(),
}),
});
// TypeScript knows the shape of parameters and output
weatherTool.toolCall("t1", {
location: "NYC",
units: "C", // ✅ Type-safe
// units: 'K' // ❌ TypeScript error
});
weatherTool.toolCallResult("t1", {
temp: 72,
condition: "Sunny", // ✅ Type-safe
// humidity: 50 // ❌ TypeScript error
});Torque executes in two phases:
This enables:
Control randomness for reproducible datasets:
await generateDataset(schema, {
count: 50,
model: openai("gpt-5-mini"),
output: "data/dataset.jsonl",
seed: 12345, // Same seed = same output
});How seeds work:
Token counts for each row are computed off the main thread using a worker pool so dataset generation stays responsive. Configure the pool with tokenCounterWorkers (default: 3), or disable counting entirely by setting it to 0.
await generateDataset(schema, {
count: 20,
model: openai("gpt-5-mini"),
tokenCounterWorkers: 5, // spawn 5 token-counting workers
});Choose your preferred output file format and data structure:
// Export as JSONL with default ai-sdk structure (default)
await generateDataset(schema, {
count: 100,
model: openai("gpt-4o-mini"),
format: "jsonl",
output: "data/dataset.jsonl",
});
// Export in OpenAI Chat Completions format (tools + messages structure)
await generateDataset(schema, {
count: 100,
model: openai("gpt-4o-mini"),
format: "jsonl",
exportFormat: "chat_template",
output: "data/finetune.jsonl",
});Supported File Formats (format):
Supported Data Structures (exportFormat):
Both formats write rows incrementally as they're generated, so large datasets won't consume excessive memory.
💡 When format is specified without output, the file extension is automatically set based on the format.
💡 See full example: examples/parquet-export.ts
Model conversations where tools take time to execute:
import {
generateDataset,
generatedUser,
generatedAssistant,
generatedToolCall,
generatedToolCallResult,
tool,
times,
between,
} from "@qforge/torque";
import { z } from "zod";
const searchTool = tool({
name: "web_search",
description: "Search the web",
parameters: z.object({ query: z.string() }),
output: z.object({ results: z.array(z.string()) }),
});
await generateDataset(
() => [
searchTool.toolFunction(),
// Initial request
generatedUser({ prompt: "Ask for information requiring web search" }),
// Tool call generated based on the user request
generatedToolCall(searchTool, "search-1"),
// Immediate acknowledgment
searchTool.toolCallResult("search-1", "<tool_ack />"),
generatedAssistant({
prompt: "Acknowledge search started, assure user it's in progress",
}),
// Filler conversation while waiting.
// While generating AI is aware how many messages are left.
times(between(1, 3), [
generatedUser({ prompt: "Casual conversation, unrelated to search" }),
generatedAssistant({ prompt: "Respond naturally to casual topic" }),
]),
// Actual result arrives with reused arguments
generatedToolCall(searchTool, "search-1-FINAL", {
reuseArgsFrom: "search-1",
}),
// Generated actual result based on previously generated tool call
generatedToolCallResult(searchTool, "search-1-FINAL"),
generatedAssistant({ prompt: "Present search results to user" }),
],
{
count: 50,
model: openai("gpt-5-mini"),
output: "data/async-tools.jsonl",
}
);💡 See full example: examples/async-tools.ts | ▶️ Try in Browser
Guide the AI's generation style globally:
await generateDataset(schema, {
count: 100,
model: openai("gpt-5-mini"),
output: "data/dataset.jsonl",
generationContext: {
global: {
messages: [
{
role: "system",
content:
'Keep messages concise and natural. Avoid starting with "Sure" or "Thanks".',
},
],
},
user: {
messages: [
{
role: "system",
content:
"Generate diverse user messages with varying levels of technical detail.",
},
],
},
assistant: {
messages: [
{
role: "system",
content:
"Assistant should be helpful but concise. Use 2-3 sentences max.",
},
],
},
},
});💡 See full example: examples/custom-generation-context.ts | ▶️ Try in Browser
Generate datasets with different tools:
import { oneOf } from "@qforge/torque";
const tools = [weatherTool, calculatorTool, searchTool];
await generateDataset(
() => {
const tool = oneOf(tools);
return [
tool.toolFunction(),
generatedUser({ prompt: "Ask question requiring this tool" }),
generatedToolCall(tool, "t1"),
generatedToolCallResult(tool, "t1"),
generatedAssistant({ prompt: "Present the result" }),
];
},
{
count: 300, // 100 examples per tool
model: openai("gpt-5-mini"),
output: "data/multi-tool.jsonl",
}
);💡 See full example: examples/multiple-tool-variations.ts | ▶️ Try in Browser
Torque includes built-in Faker.js integration that automatically respects the seed system for reproducible fake data generation:
import {
generateDataset,
generatedUser,
generatedAssistant,
faker,
} from "@qforge/torque";
await generateDataset(
() => [
generatedUser({
prompt: `Introduce yourself as ${faker.person.fullName()} from ${faker.location.city()}`,
}),
generatedAssistant({
prompt: "Greet the user warmly",
}),
],
{
count: 100,
model: openai("gpt-5-mini"),
output: "data/personas.jsonl",
seed: 42, // Same seed = same fake names and cities
}
);Faker automatically uses Torque's seed system, so:
Common use cases:
💡 See full example: examples/faker-integration.ts | ▶️ Try in Browser
Torque includes a beautiful CLI interface with:
╭────────────────────────────────────────────────────╮ │ Dataset Generation │ ├────────────────────────────────────────────────────┤ │ Total: 100 │ │ Completed: 45 │ │ In Progress: 5 │ │ Seed: 42 │ │ Output: data/dataset_2025-10-30.jsonl │ │ Workers: 5 │ ├────────────────────────────────────────────────────┤ │ ████████████░░░░░░░░░░░░░ 45% │ ├────────────────────────────────────────────────────┤ │ #0: [████████████████░░░░] 80% tool-result (search)│ │ #1: [██████░░░░░░░░░░░░░░] 30% user message │ │ #2: [████████████████████] 100% Writing... │ │ #3: [██░░░░░░░░░░░░░░░░░░] 10% assistant message │ │ #4: [██████████░░░░░░░░░░] 50% tool-call (calc) │ ╰────────────────────────────────────────────────────╯
Contributions are welcome! This is part of a larger project exploring async tool patterns in LLMs.
MIT License - see LICENSE for details
Built with:
Made with ❤️ for the AI tinkerers community
| Back | FazBrowse Home | New Git URL |