> ## Documentation Index
> Fetch the complete documentation index at: https://docs.weventures.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# AI & LLM Providers

> Language models, AI frameworks, and orchestration tools

## Major LLM Providers

<Tabs>
  <Tab title="OpenAI">
    ### OpenAI

    **GPT-5, GPT-5 Mini, GPT-5 Nano, o3, DALL-E, Whisper**

    Industry leader in large language models. GPT-5 for coding & agents, o3 for complex reasoning.

    **Best for:** Production AI apps, coding, agentic tasks, embeddings

    **Models (2025):**

    * **GPT-5** - State-of-the-art coding & agents (74.9% SWE-bench, 88% Aider)
    * **GPT-5 Mini** - Balanced performance & cost
    * **GPT-5 Nano** - Ultra-fast & affordable
    * **GPT-5 Pro** - Highest quality with scaled reasoning
    * **o3** - Advanced reasoning model
    * **DALL-E 3** - Image generation
    * **Whisper** - Speech-to-text

    **Context:** 272K input tokens, 128K output tokens

    **Pricing (per 1M tokens):**

    * GPT-5: $1.25 input / $10 output
    * GPT-5 Mini: $0.25 input / $2 output
    * GPT-5 Nano: $0.05 input / $0.40 output
    * Prompt caching: \$0.125/1M tokens

    **Key Features:**

    * 45% fewer factual errors with web search
    * 80% fewer hallucinations vs o3
    * 4 reasoning levels: minimal, low, medium, high
    * Parallel & sequential tool calling
    * Available to free users

    [Documentation](https://platform.openai.com/docs) • [API](https://openai.com/api)
  </Tab>

  <Tab title="Anthropic">
    ### Anthropic

    **Claude Sonnet 4.5, Opus 4.1, Haiku 4**

    Best coding model in the world. Long context (1M tokens), autonomous agents, safety-focused.

    **Best for:** Coding, long-running autonomous tasks, safety-critical apps

    **Models (2025):**

    * **Claude Sonnet 4.5** - Best coding model (77.2% SWE-bench, 61.4% OSWorld)
    * **Claude Opus 4.1** - Most capable for complex reasoning
    * **Claude Haiku 4** - Fastest & most affordable

    **Context:** Up to 1M tokens (with beta header)

    **Pricing (per 1M tokens):**

    * Sonnet 4.5: $3 input / $15 output
    * Opus 4.1: $15 input / $75 output
    * 90% savings with prompt caching
    * 50% savings with batch processing

    **Key Features:**

    * 64K output tokens for rich code generation
    * Runs autonomously for 30+ hours
    * State-of-the-art cybersecurity & compliance
    * More resistant to prompt injection
    * Outperforms Opus 4.1 in most areas

    [Documentation](https://docs.anthropic.com) • [API](https://anthropic.com)
  </Tab>

  <Tab title="Google AI">
    ### Google AI

    **Gemini 2.5 Pro, Flash**

    Advanced reasoning model with thinking capabilities. Native audio, 1M+ context.

    **Best for:** Complex reasoning, multimodal apps, cost-effectiveness

    **Models (2025):**

    * **Gemini 2.5 Pro** - State-of-the-art reasoning (18.8% Humanity's Last Exam)
    * **Gemini 2.5 Pro Deep Think** - Extended reasoning variant
    * **Gemini 2.0 Flash** - Ultra-fast & affordable

    **Context:** 1M tokens (2M coming soon)

    **Pricing (per 1M tokens):**

    * Standard (\<200K): $1.25 input / $10 output
    * Long context (>200K): $2.50 input / $15 output
    * Flash: $0.10 input / $0.40 output

    **Key Features:**

    * Native audio outputs (24 languages)
    * Thinking tokens for reasoning
    * 63.8% SWE-bench Verified
    * \#1 on LMArena leaderboard
    * Cheaper than Anthropic, competitive with OpenAI

    [Documentation](https://ai.google.dev/docs) • [API](https://ai.google.dev)
  </Tab>

  <Tab title="Cohere">
    ### Cohere

    **Command, Embed, Rerank**

    Enterprise-focused LLMs with strong embeddings and RAG optimization.

    **Best for:** Enterprise deployments, embeddings, retrieval-augmented generation

    **Models:**

    * **Command R+** - Most capable
    * **Command R** - Balanced performance
    * **Embed v3** - Best-in-class embeddings
    * **Rerank** - Search result ranking

    **Pricing:** \$\$\$ (enterprise pricing)

    [Documentation](https://docs.cohere.com) • [API](https://cohere.com)
  </Tab>
</Tabs>

## Fast Inference Providers

<CardGroup cols={2}>
  <Card title="Groq" icon="microchip">
    **Ultra-fast LLM inference with LPU**

    500+ tokens/sec, lowest latency in industry. Perfect for real-time applications.

    **Speed:** Fastest (LPU technology)

    **Models:** Llama 3.1, Mixtral, Gemma

    **Best for:** Real-time AI, chatbots, low-latency requirements

    **Pricing:** \$ (generous free tier)

    [Get Started](https://groq.com)
  </Card>

  <Card title="Together AI" icon="layer-group">
    **Fast inference + fine-tuning**

    Open-source models with custom fine-tuning capabilities.

    **Best for:** Custom models, fine-tuning, open-source LLMs

    **Features:**

    * Fine-tuning support
    * Open-source models
    * Competitive pricing
    * Fast inference

    **Pricing:** \$\$ (mid-range)

    [Get Started](https://together.ai)
  </Card>

  <Card title="Fireworks AI" icon="fire">
    **Production-grade inference**

    Fast inference for open models with function calling support.

    **Best for:** Serverless AI, function calling, production scale

    **Features:**

    * Serverless deployment
    * Function calling
    * Open-source models
    * Fast performance

    **Pricing:** \$\$ (pay-per-use)

    [Get Started](https://fireworks.ai)
  </Card>
</CardGroup>

## Aggregators & Multi-Model Platforms

<CardGroup cols={2}>
  <Card title="OpenRouter" icon="route">
    **Unified API for 300+ models**

    Access OpenAI, Anthropic, Google, Groq, Meta through one API with automatic fallbacks.

    **Best for:** Multi-model apps, cost optimization, avoiding vendor lock-in

    **Features:**

    * 300+ models
    * Automatic fallbacks
    * Cost optimization
    * Single API

    **Pricing:** 5% markup on base costs

    [Get Started](https://openrouter.ai)
  </Card>

  <Card title="Replicate" icon="copy">
    **Run open-source models**

    Stable Diffusion, LLaMA, Whisper, and more. Pay-per-use serverless.

    **Best for:** Open-source models, image/video generation, experimentation

    **Features:**

    * Open-source models
    * Image generation
    * Video generation
    * Serverless deployment

    **Pricing:** \$ (pay-per-use)

    [Get Started](https://replicate.com)
  </Card>
</CardGroup>

## AI Frameworks & Orchestration

<CardGroup cols={2}>
  <Card title="LangChain" icon="link">
    **Framework for building LLM apps**

    v1.0 alpha released! New create\_agent built on LangGraph runtime.

    **Best for:** RAG, chatbots, complex LLM workflows

    **Latest Features (2025):**

    * LangChain 1.0 alpha with unified agent
    * Built on LangGraph runtime
    * Python & JavaScript support
    * Enhanced observability in LangSmith
    * Open Agent Platform (no-code builder)

    [Documentation](https://python.langchain.com)
  </Card>

  <Card title="LangGraph" icon="diagram-project">
    **Build stateful AI agents (GA)**

    Production-ready platform for deploying long-running autonomous agents.

    **Best for:** Complex agents, multi-step workflows, production deployments

    **Latest Features (2025):**

    * LangGraph Platform (Generally Available)
    * Node caching & deferred nodes
    * Revision queueing for smooth deploys
    * Trace mode with LangSmith integration
    * Dynamic tool calling
    * Pre/Post model hooks
    * Built-in web search & RemoteMCP
    * Studio v2 (runs locally)
    * 1-click GitHub deployment

    [Documentation](https://langchain-ai.github.io/langgraph)
  </Card>

  <Card title="CrewAI" icon="users-gear">
    **Multi-agent orchestration**

    Role-based agents collaborating on tasks. Hierarchical or sequential workflows.

    **Best for:** Multi-agent systems, collaborative AI, task delegation

    * Multi-agent
    * Role-based
    * Collaborative workflows
    * Easy setup

    [Documentation](https://docs.crewai.com)
  </Card>

  <Card title="Vercel AI SDK" icon="triangle">
    **TypeScript AI framework**

    Streaming responses, function calling, React hooks. Built for Next.js.

    **Best for:** Next.js apps, edge AI, streaming UIs, React integration

    * TypeScript-first
    * Streaming support
    * React hooks
    * Edge compatible

    [Documentation](https://sdk.vercel.ai/docs)
  </Card>

  <Card title="LlamaIndex" icon="list">
    **Data framework for LLMs**

    Data connectors, indexes, query engines. Optimized for RAG.

    **Best for:** RAG applications, data ingestion, structured retrieval

    * Data connectors
    * RAG-optimized
    * Query engines
    * Indexing strategies

    [Documentation](https://docs.llamaindex.ai)
  </Card>

  <Card title="AutoGen" icon="robot">
    **Multi-agent framework (Microsoft)**

    Conversable agents with human-in-the-loop and code execution.

    **Best for:** Research, multi-agent collaboration, code generation

    * Multi-agent
    * Code execution
    * Human-in-loop
    * Microsoft Research

    [Documentation](https://microsoft.github.io/autogen)
  </Card>
</CardGroup>

## Provider Comparison (2025)

| Provider                  | Speed     | Cost   | Context | Best For                              |
| ------------------------- | --------- | ------ | ------- | ------------------------------------- |
| **OpenAI GPT-5**          | Fast      | \$\$   | 272K    | Coding, agents, production AI         |
| **Anthropic Sonnet 4.5**  | Medium    | \$\$\$ | 1M      | Best coding, autonomous agents        |
| **Google Gemini 2.5 Pro** | Medium    | \$\$   | 1M-2M   | Reasoning, multimodal, cost-effective |
| **Groq**                  | ⚡ Fastest | \$     | Varies  | Real-time, low latency                |
| **OpenRouter**            | Fast      | \$\$   | Varies  | Multi-model flexibility               |
| **Together AI**           | Fast      | \$\$   | Varies  | Open models, fine-tuning              |

## Pricing Comparison (per 1M tokens - 2025)

**Input Tokens:**

* GPT-5: \$1.25
* GPT-5 Mini: \$0.25
* GPT-5 Nano: \$0.05
* Claude Sonnet 4.5: \$3.00
* Claude Opus 4.1: \$15.00
* Gemini 2.5 Pro (standard): \$1.25
* Gemini 2.5 Pro (long context): \$2.50
* Gemini 2.0 Flash: \$0.10

**Output Tokens:**

* GPT-5: \$10.00
* GPT-5 Mini: \$2.00
* GPT-5 Nano: \$0.40
* Claude Sonnet 4.5: \$15.00
* Claude Opus 4.1: \$75.00
* Gemini 2.5 Pro (standard): \$10.00
* Gemini 2.5 Pro (long context): \$15.00
* Gemini 2.0 Flash: \$0.40

## Edge Functions vs Custom SDKs

<Tabs>
  <Tab title="Supabase Edge Functions">
    **Deno-based serverless functions**

    **Pros:**

    * Direct database access
    * Global deployment
    * TypeScript/JavaScript
    * Built-in secrets management
    * Free tier included

    **Cons:**

    * Deno runtime only
    * Supabase ecosystem

    **Best for:** Supabase users, TypeScript developers, database-connected AI

    ```typescript theme={null}
    // Supabase Edge Function example
    import { serve } from "https://deno.land/std@0.168.0/http/server.ts"

    serve(async (req) => {
      const { prompt } = await req.json()

      const response = await fetch("https://api.openai.com/v1/chat/completions", {
        method: "POST",
        headers: {
          "Authorization": `Bearer ${Deno.env.get("OPENAI_API_KEY")}`,
          "Content-Type": "application/json"
        },
        body: JSON.stringify({
          model: "gpt-4",
          messages: [{ role: "user", content: prompt }]
        })
      })

      return new Response(JSON.stringify(await response.json()))
    })
    ```
  </Tab>

  <Tab title="Provider SDKs">
    **Official OpenAI, Anthropic SDKs**

    **Pros:**

    * Official support
    * Latest features
    * Type-safe
    * Well-documented
    * Works anywhere

    **Cons:**

    * More boilerplate
    * Need separate hosting
    * API key management

    **Best for:** Any runtime, maximum flexibility, latest features

    ```typescript theme={null}
    // OpenAI SDK example
    import OpenAI from "openai"

    const openai = new OpenAI({
      apiKey: process.env.OPENAI_API_KEY
    })

    const completion = await openai.chat.completions.create({
      model: "gpt-4",
      messages: [{ role: "user", content: "Hello!" }]
    })
    ```
  </Tab>
</Tabs>

## Memory & State Management

<CardGroup cols={2}>
  <Card title="Zustand" icon="box">
    **Lightweight React state**

    Perfect for managing AI conversation state in React apps.

    * Minimal boilerplate
    * React integration
    * TypeScript support
    * Persistent storage

    [Documentation](https://zustand-demo.pmnd.rs)
  </Card>

  <Card title="LangChain Memory" icon="brain">
    **Built-in conversation memory**

    Buffer, summary, entity, and vector store memory types.

    * Conversation history
    * Summary memory
    * Entity extraction
    * Vector store integration

    [Documentation](https://python.langchain.com/docs/modules/memory)
  </Card>

  <Card title="Upstash Redis" icon="database">
    **Serverless Redis for state**

    Perfect for storing conversation history and session data.

    * Serverless
    * Low latency
    * Global replication
    * REST API

    [Documentation](https://upstash.com/docs/redis)
  </Card>

  <Card title="Vercel KV" icon="key">
    **Edge-compatible key-value store**

    Built on Upstash, perfect for Next.js edge functions.

    * Edge-compatible
    * Low latency
    * Simple API
    * Vercel integration

    [Documentation](https://vercel.com/docs/storage/vercel-kv)
  </Card>
</CardGroup>

## Decision Trees

<Accordion title="Choosing an LLM Provider">
  **Need industry-leading performance?** → OpenAI (GPT-4)

  **Long documents & safety-critical?** → Anthropic (Claude)

  **Real-time, low-latency chatbots?** → Groq

  **Want flexibility across providers?** → OpenRouter

  **Open-source models & fine-tuning?** → Together AI

  **Image/video generation?** → Replicate

  **Cost-effectiveness & multimodal?** → Google AI (Gemini)
</Accordion>

<Accordion title="Choosing an AI Framework">
  **Building RAG applications?** → LangChain + LlamaIndex

  **Complex multi-step agents?** → LangGraph

  **Multi-agent collaboration?** → CrewAI

  **Next.js with streaming UI?** → Vercel AI SDK

  **Research & experimentation?** → AutoGen

  **Simple chat integration?** → Direct SDK (OpenAI, Anthropic)
</Accordion>

## Recommended Combinations

<Steps>
  <Step title="Production RAG App">
    **OpenAI** (embeddings) + **Pinecone** (vectors) + **LangChain** (orchestration)
  </Step>

  <Step title="Real-time Chatbot">
    **Groq** (inference) + **Upstash Redis** (state) + **Vercel AI SDK** (streaming)
  </Step>

  <Step title="Cost-Optimized">
    **OpenRouter** (multi-model) + **Chroma** (vectors) + **Supabase Edge Functions**
  </Step>

  <Step title="Enterprise AI">
    **Anthropic Claude** (safety) + **Weaviate** (hybrid search) + **LangGraph** (agents)
  </Step>
</Steps>
