AsyncClient / SyncClient

BAML generates both a sync client and an async client. They offer the exact same public API but methods are either synchronous or asynchronous.

BAML Functions

The generated client exposes all the functions that you’ve defined your BAML files as methods. Suppose we have this file named baml_src/literature.baml:

baml_src/literature.baml
function TellMeAStory() -> string {
client "openai/gpt-4o"
prompt #"
Tell me a story
"#
}
function WriteAPoemAbout(input: string) -> string {
client "openai/gpt-4o"
prompt #"
Write a poem about {{ input }}
"#
}

After running baml-cli generate you can directly call these functions from your code. Here’s an example using the async client:

from baml_client.async_client import b
async def example():
# Call your BAML functions.
story = await b.TellMeAStory()
poem = await b.WriteAPoemAbout("Roses")

The sync client is exactly the same but it doesn’t need an async runtime, instead it just blocks.

from baml_client.sync_client import b
def example():
# Call your BAML functions.
story = b.TellMeAStory()
poem = b.WriteAPoemAbout("Roses")

Call Patterns

The client object exposes some references to other objects that call your functions in a different manner.

.stream

The .stream object is used to stream the response from a function.

from baml_client.async_client import b
async def example():
stream = b.stream.TellMeAStory()
async for partial in stream:
print(partial)
print(await stream.get_final_response())

.request

This feature was added in: v0.79.0

The .request object returns the raw HTTP request but it does not send it. However, the async client still returns an awaitable object because we might need to resolve media types like images and convert them to base64 or the required format in order to send them to the LLM.

from baml_client.async_client import b
async def example():
request = await b.request.TellMeAStory()
print(request.url)
print(request.headers)
print(request.body.json())

.stream_request

This feature was added in: v0.79.0

Same as .request but sets the streaming options to true.

from baml_client.async_client import b
async def example():
request = await b.stream_request.TellMeAStory()
print(request.url)
print(request.headers)
print(request.body.json())

.parse

This feature was added in: v0.79.0

The .parse object is used to parse the response returned by the LLM after the function call. Can be used in combination with .request.

import requests
# requests is not async so for simplicity we'll use the sync client.
from baml_client.sync_client import b
def example():
# Get the HTTP request.
request = b.request.TellMeAStory()
# Send the HTTP request.
response = requests.post(request.url, headers=request.headers, json=request.body.json())
# Parse the LLM response.
parsed = b.parse.TellMeAStory(response.json()["choices"][0]["message"]["content"])
# Fully parsed response.
print(parsed)

.parse_stream

This feature was added in: v0.79.0

Same as .parse but for streaming responses. Can be used in combination with .stream_request.

from openai import AsyncOpenAI
from baml_client.async_client import b
async def example():
client = AsyncOpenAI()
request = await b.stream_request.TellMeAStory()
stream = await client.chat.completions.create(**request.body.json())
llm_response: list[str] = []
async for chunk in stream:
if len(chunk.choices) > 0 and chunk.choices[0].delta.content is not None:
llm_response.append(chunk.choices[0].delta.content)
print(b.parse_stream.TellMeAStory("".join(llm_response)))