📎 feat: Direct Provider Attachment Support for Multimodal Content (#9994)

* 📎 feat: Direct Provider Attachment Support for Multimodal Content * 📑 feat: Anthropic Direct Provider Upload (#9072) * feat: implement Anthropic native PDF support with document preservation - Add comprehensive debug logging throughout PDF processing pipeline - Refactor attachment processing to separate image and document handling - Create distinct addImageURLs(), addDocuments(), and processAttachments() methods - Fix critical bugs in stream handling and parameter passing - Add streamToBuffer utility for proper stream-to-buffer conversion - Remove api/agents submodule from repository 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * chore: remove out of scope formatting changes * fix: stop duplication of file in chat on end of response stream * chore: bring back file search and ocr options * chore: localize upload to provider string in file menu * refactor: change createMenuItems args to fit new pattern introduced by anthropic-native-pdf-support * feat: add cache point for pdfs processed by anthropic endpoint since they are unlikely to change and should benefit from caching * feat: combine Upload Image into Upload to Provider since they both perform direct upload and change provider upload icon to reflect multimodal upload * feat: add citations support according to docs * refactor: remove redundant 'document' check since documents are handled properly by formatMessage in the agents repo now * refactor: change upload logic so anthropic endpoint isn't exempted from normal upload path using Agents for consistency with the rest of the upload logic * fix: include width and height in return from uploadLocalFile so images are correctly identified when going through an AgentUpload in addImageURLs * chore: remove client specific handling since the direct provider stuff is handled by the agent client * feat: handle documents in AgentClient so no need for change to agents repo * chore: removed unused changes * chore: remove auto generated comments from OG commit * feat: add logic for agents to use direct to provider uploads if supported (currently just anthropic) * fix: reintroduce role check to fix render error because of undefined value for Content Part * fix: actually fix render bug by using proper isCreatedByUser check and making sure our mutation of formattedMessage.content is consistent --------- Co-authored-by: Andres Restrepo <andres@thelinuxkid.com> Co-authored-by: Claude <noreply@anthropic.com> 📁 feat: Send Attachments Directly to Provider (OpenAI) (#9098) * refactor: change references from direct upload to direct attach to better reflect functionality since we are just using base64 encoding strategy now rather than Files/File API for sending our attachments directly to the provider, the upload nomenclature no longer makes sense. direct_attach better describes the different methods of sending attachments to providers anyways even if we later introduce direct upload support * feat: add upload to provider option for openai (and agent) ui * chore: move anthropic pdf validator over to packages/api * feat: simple pdf validation according to openai docs * feat: add provider agnostic validatePdf logic to start handling multiple endpoints * feat: add handling for openai specific documentPart formatting * refactor: move require statement to proper place at top of file * chore: add in openAI endpoint for the rest of the document handling logic * feat: add direct attach support for azureOpenAI endpoint and agents * feat: add pdf validation for azureOpenAI endpoint * refactor: unify all the endpoint checks with isDocumentSupportedEndpoint * refactor: consolidate Upload to Provider vs Upload image logic for clarity * refactor: remove anthropic from anthropic_multimodal fileType since we support multiple providers now 🗂️ feat: Send Attachments Directly to Provider (Google) (#9100) * feat: add validation for google PDFs and add google endpoint as a document supporting endpoint * feat: add proper pdf formatting for google endpoints (requires PR #14 in agents) * feat: add multimodal support for google endpoint attachments * feat: add audio file svg * fix: refactor attachments logic so multi-attachment messages work properly * feat: add video file svg * fix: allows for followup questions of uploaded multimodal attachments * fix: remove incorrect final message filtering that was breaking Attachment component rendering fix: manualy rename 'documents' to 'Documents' in git since it wasn't picked up due to case insensitivity in dir name fix: add logic so filepicker for a google agent has proper filetype filtering 🛫 refactor: Move Encoding Logic to packages/api (#9182) * refactor: move audio encode over to TS * refactor: audio encoding now functional in LC again * refactor: move video encode over to TS * refactor: move document encode over to TS * refactor: video encoding now functional in LC again * refactor: document encoding now functional in LC again * fix: extend file type options in AttachFileMenu to include 'google_multimodal' and update dependency array to include agent?.provider * feat: only accept pdfs if responses api is enabled for openai convos chore: address ESLint comments chore: add missing audio mimetype * fix: type safety for message content parts and improve null handling * chore: reorder AttachFileMenuProps for consistency and clarity * chore: import order in AttachFileMenu * fix: improve null handling for text parts in parseTextParts function * fix: remove no longer used unsupported capability error message for file uploads * fix: OpenAI Direct File Attachment Format * fix: update encodeAndFormatDocuments to support OpenAI responses API and enhance document result types * refactor: broaden providers supported for documents * feat: enhance DragDrop context and modal to support document uploads based on provider capabilities * fix: reorder import statements for consistency in video encoding module --------- Co-authored-by: Dustin Healy <54083382+dustinhealy@users.noreply.github.com>
2026-02-04 00:31:50 +01:00 · 2025-10-06 17:30:16 -04:00 · 2025-10-06 17:30:16 -04:00 · bcd97aad2f
commit bcd97aad2f
parent 9c77f53454
33 changed files with 1040 additions and 74 deletions
--- a/packages/api/src/files/encode/audio.ts
+++ b/packages/api/src/files/encode/audio.ts
@ -0,0 +1,74 @@
+import { Providers } from '@librechat/agents';
+import { isDocumentSupportedProvider } from 'librechat-data-provider';
+import type { IMongoFile } from '@librechat/data-schemas';
+import type { Request } from 'express';
+import type { StrategyFunctions, AudioResult } from '~/types/files';
+import { validateAudio } from '~/files/validation';
+import { getFileStream } from './utils';
+
+/**
+ * Encodes and formats audio files for different providers
+ * @param req - The request object
+ * @param files - Array of audio files
+ * @param provider - The provider to format for (currently only google is supported)
+ * @param getStrategyFunctions - Function to get strategy functions
+ * @returns Promise that resolves to audio and file metadata
+ */
+export async function encodeAndFormatAudios(
+  req: Request,
+  files: IMongoFile[],
+  provider: Providers,
+  getStrategyFunctions: (source: string) => StrategyFunctions,
+): Promise<AudioResult> {
+  if (!files?.length) {
+    return { audios: [], files: [] };
+  }
+
+  const encodingMethods: Record<string, StrategyFunctions> = {};
+  const result: AudioResult = { audios: [], files: [] };
+
+  const results = await Promise.allSettled(
+    files.map((file) => getFileStream(req, file, encodingMethods, getStrategyFunctions)),
+  );
+
+  for (const settledResult of results) {
+    if (settledResult.status === 'rejected') {
+      console.error('Audio processing failed:', settledResult.reason);
+      continue;
+    }
+
+    const processed = settledResult.value;
+    if (!processed) continue;
+
+    const { file, content, metadata } = processed;
+
+    if (!content || !file) {
+      if (metadata) result.files.push(metadata);
+      continue;
+    }
+
+    if (!file.type.startsWith('audio/') || !isDocumentSupportedProvider(provider)) {
+      result.files.push(metadata);
+      continue;
+    }
+
+    const audioBuffer = Buffer.from(content, 'base64');
+    const validation = await validateAudio(audioBuffer, audioBuffer.length, provider);
+
+    if (!validation.isValid) {
+      throw new Error(`Audio validation failed: ${validation.error}`);
+    }
+
+    if (provider === Providers.GOOGLE || provider === Providers.VERTEXAI) {
+      result.audios.push({
+        type: 'audio',
+        mimeType: file.type,
+        data: content,
+      });
+    }
+
+    result.files.push(metadata);
+  }
+
+  return result;
+}
--- a/packages/api/src/files/encode/document.ts
+++ b/packages/api/src/files/encode/document.ts
@ -0,0 +1,108 @@
+import { Providers } from '@librechat/agents';
+import { isOpenAILikeProvider, isDocumentSupportedProvider } from 'librechat-data-provider';
+import type { IMongoFile } from '@librechat/data-schemas';
+import type { Request } from 'express';
+import type { StrategyFunctions, DocumentResult } from '~/types/files';
+import { validatePdf } from '~/files/validation';
+import { getFileStream } from './utils';
+
+/**
+ * Processes and encodes document files for various providers
+ * @param req - Express request object
+ * @param files - Array of file objects to process
+ * @param provider - The provider name
+ * @param getStrategyFunctions - Function to get strategy functions
+ * @returns Promise that resolves to documents and file metadata
+ */
+export async function encodeAndFormatDocuments(
+  req: Request,
+  files: IMongoFile[],
+  { provider, useResponsesApi }: { provider: Providers; useResponsesApi?: boolean },
+  getStrategyFunctions: (source: string) => StrategyFunctions,
+): Promise<DocumentResult> {
+  if (!files?.length) {
+    return { documents: [], files: [] };
+  }
+
+  const encodingMethods: Record<string, StrategyFunctions> = {};
+  const result: DocumentResult = { documents: [], files: [] };
+
+  const documentFiles = files.filter(
+    (file) => file.type === 'application/pdf' || file.type?.startsWith('application/'),
+  );
+
+  if (!documentFiles.length) {
+    return result;
+  }
+
+  const results = await Promise.allSettled(
+    documentFiles.map((file) => {
+      if (file.type !== 'application/pdf' || !isDocumentSupportedProvider(provider)) {
+        return Promise.resolve(null);
+      }
+      return getFileStream(req, file, encodingMethods, getStrategyFunctions);
+    }),
+  );
+
+  for (const settledResult of results) {
+    if (settledResult.status === 'rejected') {
+      console.error('Document processing failed:', settledResult.reason);
+      continue;
+    }
+
+    const processed = settledResult.value;
+    if (!processed) continue;
+
+    const { file, content, metadata } = processed;
+
+    if (!content || !file) {
+      if (metadata) result.files.push(metadata);
+      continue;
+    }
+
+    if (file.type === 'application/pdf' && isDocumentSupportedProvider(provider)) {
+      const pdfBuffer = Buffer.from(content, 'base64');
+      const validation = await validatePdf(pdfBuffer, pdfBuffer.length, provider);
+
+      if (!validation.isValid) {
+        throw new Error(`PDF validation failed: ${validation.error}`);
+      }
+
+      if (provider === Providers.ANTHROPIC) {
+        result.documents.push({
+          type: 'document',
+          source: {
+            type: 'base64',
+            media_type: 'application/pdf',
+            data: content,
+          },
+          cache_control: { type: 'ephemeral' },
+          citations: { enabled: true },
+        });
+      } else if (useResponsesApi) {
+        result.documents.push({
+          type: 'input_file',
+          filename: file.filename,
+          file_data: `data:application/pdf;base64,${content}`,
+        });
+      } else if (provider === Providers.GOOGLE || provider === Providers.VERTEXAI) {
+        result.documents.push({
+          type: 'document',
+          mimeType: 'application/pdf',
+          data: content,
+        });
+      } else if (isOpenAILikeProvider(provider) && provider != Providers.AZURE) {
+        result.documents.push({
+          type: 'file',
+          file: {
+            filename: file.filename,
+            file_data: `data:application/pdf;base64,${content}`,
+          },
+        });
+      }
+      result.files.push(metadata);
+    }
+  }
+
+  return result;
+}
--- a/packages/api/src/files/encode/index.ts
+++ b/packages/api/src/files/encode/index.ts
@ -0,0 +1,3 @@
+export * from './audio';
+export * from './document';
+export * from './video';
--- a/packages/api/src/files/encode/utils.ts
+++ b/packages/api/src/files/encode/utils.ts
@ -0,0 +1,46 @@
+import getStream from 'get-stream';
+import { FileSources } from 'librechat-data-provider';
+import type { IMongoFile } from '@librechat/data-schemas';
+import type { Request } from 'express';
+import type { StrategyFunctions, ProcessedFile } from '~/types/files';
+
+/**
+ * Processes a file by downloading and encoding it to base64
+ * @param req - Express request object
+ * @param file - File object to process
+ * @param encodingMethods - Cache of encoding methods by source
+ * @param getStrategyFunctions - Function to get strategy functions for a source
+ * @returns Processed file with content and metadata, or null if filepath missing
+ */
+export async function getFileStream(
+  req: Request,
+  file: IMongoFile,
+  encodingMethods: Record<string, StrategyFunctions>,
+  getStrategyFunctions: (source: string) => StrategyFunctions,
+): Promise<ProcessedFile | null> {
+  if (!file?.filepath) {
+    return null;
+  }
+
+  const source = file.source ?? FileSources.local;
+  if (!encodingMethods[source]) {
+    encodingMethods[source] = getStrategyFunctions(source);
+  }
+
+  const { getDownloadStream } = encodingMethods[source];
+  const stream = await getDownloadStream(req, file.filepath);
+  const buffer = await getStream.buffer(stream);
+
+  return {
+    file,
+    content: buffer.toString('base64'),
+    metadata: {
+      file_id: file.file_id,
+      temp_file_id: file.temp_file_id,
+      filepath: file.filepath,
+      source: file.source,
+      filename: file.filename,
+      type: file.type,
+    },
+  };
+}
--- a/packages/api/src/files/encode/video.ts
+++ b/packages/api/src/files/encode/video.ts
@ -0,0 +1,74 @@
+import { Providers } from '@librechat/agents';
+import { isDocumentSupportedProvider } from 'librechat-data-provider';
+import type { IMongoFile } from '@librechat/data-schemas';
+import type { Request } from 'express';
+import type { StrategyFunctions, VideoResult } from '~/types/files';
+import { validateVideo } from '~/files/validation';
+import { getFileStream } from './utils';
+
+/**
+ * Encodes and formats video files for different providers
+ * @param req - The request object
+ * @param files - Array of video files
+ * @param provider - The provider to format for
+ * @param getStrategyFunctions - Function to get strategy functions
+ * @returns Promise that resolves to videos and file metadata
+ */
+export async function encodeAndFormatVideos(
+  req: Request,
+  files: IMongoFile[],
+  provider: Providers,
+  getStrategyFunctions: (source: string) => StrategyFunctions,
+): Promise<VideoResult> {
+  if (!files?.length) {
+    return { videos: [], files: [] };
+  }
+
+  const encodingMethods: Record<string, StrategyFunctions> = {};
+  const result: VideoResult = { videos: [], files: [] };
+
+  const results = await Promise.allSettled(
+    files.map((file) => getFileStream(req, file, encodingMethods, getStrategyFunctions)),
+  );
+
+  for (const settledResult of results) {
+    if (settledResult.status === 'rejected') {
+      console.error('Video processing failed:', settledResult.reason);
+      continue;
+    }
+
+    const processed = settledResult.value;
+    if (!processed) continue;
+
+    const { file, content, metadata } = processed;
+
+    if (!content || !file) {
+      if (metadata) result.files.push(metadata);
+      continue;
+    }
+
+    if (!file.type.startsWith('video/') || !isDocumentSupportedProvider(provider)) {
+      result.files.push(metadata);
+      continue;
+    }
+
+    const videoBuffer = Buffer.from(content, 'base64');
+    const validation = await validateVideo(videoBuffer, videoBuffer.length, provider);
+
+    if (!validation.isValid) {
+      throw new Error(`Video validation failed: ${validation.error}`);
+    }
+
+    if (provider === Providers.GOOGLE || provider === Providers.VERTEXAI) {
+      result.videos.push({
+        type: 'video',
+        mimeType: file.type,
+        data: content,
+      });
+    }
+
+    result.files.push(metadata);
+  }
+
+  return result;
+}
--- a/packages/api/src/files/index.ts
+++ b/packages/api/src/files/index.ts
@ -1,5 +1,7 @@
 export * from './audio';
+export * from './encode';
 export * from './mistral/crud';
 export * from './ocr';
 export * from './parse';
+export * from './validation';
 export * from './text';
--- a/packages/api/src/files/validation.ts
+++ b/packages/api/src/files/validation.ts
@ -0,0 +1,186 @@
+import { Providers } from '@librechat/agents';
+import { mbToBytes, isOpenAILikeProvider } from 'librechat-data-provider';
+
+export interface PDFValidationResult {
+  isValid: boolean;
+  error?: string;
+}
+
+export interface VideoValidationResult {
+  isValid: boolean;
+  error?: string;
+}
+
+export interface AudioValidationResult {
+  isValid: boolean;
+  error?: string;
+}
+
+export async function validatePdf(
+  pdfBuffer: Buffer,
+  fileSize: number,
+  provider: Providers,
+): Promise<PDFValidationResult> {
+  if (provider === Providers.ANTHROPIC) {
+    return validateAnthropicPdf(pdfBuffer, fileSize);
+  }
+
+  if (isOpenAILikeProvider(provider)) {
+    return validateOpenAIPdf(fileSize);
+  }
+
+  if (provider === Providers.GOOGLE || provider === Providers.VERTEXAI) {
+    return validateGooglePdf(fileSize);
+  }
+
+  return { isValid: true };
+}
+
+/**
+ * Validates if a PDF meets Anthropic's requirements
+ * @param pdfBuffer - The PDF file as a buffer
+ * @param fileSize - The file size in bytes
+ * @returns Promise that resolves to validation result
+ */
+async function validateAnthropicPdf(
+  pdfBuffer: Buffer,
+  fileSize: number,
+): Promise<PDFValidationResult> {
+  try {
+    if (fileSize > mbToBytes(32)) {
+      return {
+        isValid: false,
+        error: `PDF file size (${Math.round(fileSize / (1024 * 1024))}MB) exceeds Anthropic's 32MB limit`,
+      };
+    }
+
+    if (!pdfBuffer || pdfBuffer.length < 5) {
+      return {
+        isValid: false,
+        error: 'Invalid PDF file: too small or corrupted',
+      };
+    }
+
+    const pdfHeader = pdfBuffer.subarray(0, 5).toString();
+    if (!pdfHeader.startsWith('%PDF-')) {
+      return {
+        isValid: false,
+        error: 'Invalid PDF file: missing PDF header',
+      };
+    }
+
+    const pdfContent = pdfBuffer.toString('binary');
+    if (
+      pdfContent.includes('/Encrypt ') ||
+      pdfContent.includes('/U (') ||
+      pdfContent.includes('/O (')
+    ) {
+      return {
+        isValid: false,
+        error: 'PDF is password-protected or encrypted. Anthropic requires unencrypted PDFs.',
+      };
+    }
+
+    const pageMatches = pdfContent.match(/\/Type[\s]*\/Page[^s]/g);
+    const estimatedPages = pageMatches ? pageMatches.length : 1;
+
+    if (estimatedPages > 100) {
+      return {
+        isValid: false,
+        error: `PDF has approximately ${estimatedPages} pages, exceeding Anthropic's 100-page limit`,
+      };
+    }
+
+    return { isValid: true };
+  } catch (error) {
+    console.error('PDF validation error:', error);
+    return {
+      isValid: false,
+      error: 'Failed to validate PDF file',
+    };
+  }
+}
+
+async function validateOpenAIPdf(fileSize: number): Promise<PDFValidationResult> {
+  if (fileSize > 10 * 1024 * 1024) {
+    return {
+      isValid: false,
+      error: "PDF file size exceeds OpenAI's 10MB limit",
+    };
+  }
+
+  return { isValid: true };
+}
+
+async function validateGooglePdf(fileSize: number): Promise<PDFValidationResult> {
+  if (fileSize > 20 * 1024 * 1024) {
+    return {
+      isValid: false,
+      error: "PDF file size exceeds Google's 20MB limit",
+    };
+  }
+
+  return { isValid: true };
+}
+
+/**
+ * Validates video files for different providers
+ * @param videoBuffer - The video file as a buffer
+ * @param fileSize - The file size in bytes
+ * @param provider - The provider to validate for
+ * @returns Promise that resolves to validation result
+ */
+export async function validateVideo(
+  videoBuffer: Buffer,
+  fileSize: number,
+  provider: Providers,
+): Promise<VideoValidationResult> {
+  if (provider === Providers.GOOGLE || provider === Providers.VERTEXAI) {
+    if (fileSize > 20 * 1024 * 1024) {
+      return {
+        isValid: false,
+        error: `Video file size (${Math.round(fileSize / (1024 * 1024))}MB) exceeds Google's 20MB limit`,
+      };
+    }
+  }
+
+  if (!videoBuffer || videoBuffer.length < 10) {
+    return {
+      isValid: false,
+      error: 'Invalid video file: too small or corrupted',
+    };
+  }
+
+  return { isValid: true };
+}
+
+/**
+ * Validates audio files for different providers
+ * @param audioBuffer - The audio file as a buffer
+ * @param fileSize - The file size in bytes
+ * @param provider - The provider to validate for
+ * @returns Promise that resolves to validation result
+ */
+export async function validateAudio(
+  audioBuffer: Buffer,
+  fileSize: number,
+  provider: Providers,
+): Promise<AudioValidationResult> {
+  if (provider === Providers.GOOGLE || provider === Providers.VERTEXAI) {
+    if (fileSize > 20 * 1024 * 1024) {
+      return {
+        isValid: false,
+        error: `Audio file size (${Math.round(fileSize / (1024 * 1024))}MB) exceeds Google's 20MB limit`,
+      };
+    }
+  }
+
+  if (!audioBuffer || audioBuffer.length < 10) {
+    return {
+      isValid: false,
+      error: 'Invalid audio file: too small or corrupted',
+    };
+  }
+
+  return { isValid: true };
+}
--- a/packages/api/src/types/files.ts
+++ b/packages/api/src/types/files.ts
@ -1,4 +1,7 @@
+import type { IMongoFile } from '@librechat/data-schemas';
 import type { ServerRequest } from './http';
+import type { Readable } from 'stream';
+import type { Request } from 'express';
 export interface STTService {
  getInstance(): Promise<STTService>;
  getProviderSchema(req: ServerRequest): Promise<[string, object]>;
@ -26,3 +29,85 @@ export interface AudioProcessingResult {
  text: string;
  bytes: number;
 }
+
+export interface VideoResult {
+  videos: Array<{
+    type: string;
+    mimeType: string;
+    data: string;
+  }>;
+  files: Array<{
+    file_id?: string;
+    temp_file_id?: string;
+    filepath: string;
+    source?: string;
+    filename: string;
+    type: string;
+  }>;
+}
+
+export interface DocumentResult {
+  documents: Array<{
+    type: 'document' | 'file' | 'input_file';
+    /** Anthropic File Format, `document` */
+    source?: {
+      type: string;
+      media_type: string;
+      data: string;
+    };
+    cache_control?: { type: string };
+    citations?: { enabled: boolean };
+    /** Google File Format, `document` */
+    mimeType?: string;
+    data?: string;
+    /** OpenAI File Format, `file` */
+    file?: {
+      filename?: string;
+      file_data?: string;
+    };
+    /** OpenAI Responses API File Format, `input_file` */
+    filename?: string;
+    file_data?: string;
+  }>;
+  files: Array<{
+    file_id?: string;
+    temp_file_id?: string;
+    filepath: string;
+    source?: string;
+    filename: string;
+    type: string;
+  }>;
+}
+
+export interface AudioResult {
+  audios: Array<{
+    type: string;
+    mimeType: string;
+    data: string;
+  }>;
+  files: Array<{
+    file_id?: string;
+    temp_file_id?: string;
+    filepath: string;
+    source?: string;
+    filename: string;
+    type: string;
+  }>;
+}
+
+export interface ProcessedFile {
+  file: IMongoFile;
+  content: string;
+  metadata: {
+    file_id: string;
+    temp_file_id?: string;
+    filepath: string;
+    source?: string;
+    filename: string;
+    type: string;
+  };
+}
+
+export interface StrategyFunctions {
+  getDownloadStream: (req: Request, filepath: string) => Promise<Readable>;
+}