JoyCaption API for Python
Python client for a JoyCaption-style image captioning API: describe images through a hosted vision-language endpoint
Run JoyCaption in the browser → View on GitHub
Install
pip install git+https://github.com/joycaption/joycaption-api.git
export SYNEXA_API_KEY="sk-..."
Quickstart
import joycaption_api
output = joycaption_api.run({
"image": "https://example.com/input.png",
"prompt": "Are you allowed to swim here?"
})
print(output)
Hosted models
- yorickvp/llava-13b — Visual instruction tuning towards large language and vision models with GPT-4 level capabilities ($0.0005 / run)
About JoyCaption
JoyCaption is an open image captioning model by fpgaminer that writes dense, literal descriptions and tag lists for diffusion-model training data, pairing a SigLIP vision encoder with a Llama 3.1 8B language model. This client calls a hosted vision-language endpoint (yorickvp/llava-13b, serving LLaVA 1.5 13B) from Python with one dependency: send an image URL and a prompt, get text back at $0.0005 per run. JoyCaption's own weights can be self-hosted from the official repository.
FAQ
Is there a JoyCaption API?
There is no official hosted JoyCaption API; the project ships weights for self-hosting. This client gives you the same image-to-text capability through a hosted VLM endpoint (yorickvp/llava-13b) that you call over HTTPS.
How much does the JoyCaption API cost?
The hosted yorickvp/llava-13b model is $0.0005 per run. Billing is per prediction; there is no hourly GPU charge, so 1,000 captions cost roughly $0.50.
Can I run JoyCaption without a GPU?
With this client, yes: inference runs on the hosted service and your code only makes HTTP requests. Self-hosting JoyCaption itself needs a CUDA GPU with enough memory for an 8B language model plus a vision encoder.
Does this client work with the original JoyCaption repo or ComfyUI?
No. It does not load the fpgaminer/joycaption checkpoints and it is not a ComfyUI node. It is a network client for the hosted endpoint. To use JoyCaption's exact modes and weights, run the official repository locally.
What input formats does it accept?
Two required fields: image (a publicly reachable image URL) and prompt (the instruction, e.g. a caption request or a question). Optional max_tokens, temperature and top_p control length and randomness. The output is plain text.
Is this the official JoyCaption SDK?
No. This is an independent, community-maintained client and is not affiliated with the JoyCaption author. The official project lives at https://github.com/fpgaminer/joycaption.