640 lines
28 KiB
HTML
640 lines
28 KiB
HTML
<!DOCTYPE html>
|
||
<html lang="fa" dir="rtl">
|
||
<head>
|
||
<meta charset="UTF-8">
|
||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||
<title>طراحی سیستم استخراج محتوا از فایل با مدل Qwen3.5 GGUF - فن آوران ساحر علم</title>
|
||
<style>
|
||
:root {
|
||
--primary: #1e3a8a;
|
||
--secondary: #3b82f6;
|
||
--accent: #f59e0b;
|
||
--success: #10b981;
|
||
--danger: #ef4444;
|
||
--warning: #f97316;
|
||
--bg-light: #f8fafc;
|
||
--bg-code: #1e293b;
|
||
--text-dark: #0f172a;
|
||
--text-muted: #64748b;
|
||
--border: #e2e8f0;
|
||
}
|
||
* { box-sizing: border-box; margin: 0; padding: 0; }
|
||
body {
|
||
font-family: 'Tahoma', 'Segoe UI', sans-serif;
|
||
background: linear-gradient(135deg, #f8fafc 0%, #e0e7ff 100%);
|
||
color: var(--text-dark);
|
||
line-height: 1.8;
|
||
padding: 20px;
|
||
}
|
||
.container {
|
||
max-width: 1200px;
|
||
margin: 0 auto;
|
||
background: white;
|
||
border-radius: 16px;
|
||
box-shadow: 0 20px 60px rgba(0,0,0,0.1);
|
||
overflow: hidden;
|
||
}
|
||
.header {
|
||
background: linear-gradient(135deg, var(--primary) 0%, var(--secondary) 100%);
|
||
color: white;
|
||
padding: 40px;
|
||
text-align: center;
|
||
position: relative;
|
||
}
|
||
.header h1 { font-size: 2.2em; margin-bottom: 10px; position: relative; }
|
||
.header .subtitle { font-size: 1.1em; opacity: 0.95; position: relative; }
|
||
.meta-bar {
|
||
display: flex;
|
||
justify-content: space-between;
|
||
background: var(--bg-light);
|
||
padding: 15px 30px;
|
||
border-bottom: 2px solid var(--border);
|
||
flex-wrap: wrap;
|
||
gap: 15px;
|
||
}
|
||
.meta-item { display: flex; align-items: center; gap: 8px; font-size: 0.9em; color: var(--text-muted); }
|
||
.meta-item strong { color: var(--primary); }
|
||
.content { padding: 40px; }
|
||
.section {
|
||
margin-bottom: 35px;
|
||
padding: 25px;
|
||
background: var(--bg-light);
|
||
border-radius: 12px;
|
||
border-right: 5px solid var(--secondary);
|
||
}
|
||
.section h2 {
|
||
color: var(--primary);
|
||
font-size: 1.6em;
|
||
margin-bottom: 20px;
|
||
padding-bottom: 10px;
|
||
border-bottom: 2px solid var(--border);
|
||
display: flex;
|
||
align-items: center;
|
||
gap: 10px;
|
||
}
|
||
.section h3 { color: var(--secondary); font-size: 1.2em; margin: 20px 0 12px; }
|
||
.step {
|
||
background: white;
|
||
padding: 20px;
|
||
margin: 15px 0;
|
||
border-radius: 10px;
|
||
box-shadow: 0 2px 8px rgba(0,0,0,0.05);
|
||
border-right: 4px solid var(--accent);
|
||
}
|
||
.step-number {
|
||
display: inline-block;
|
||
background: var(--accent);
|
||
color: white;
|
||
width: 32px;
|
||
height: 32px;
|
||
border-radius: 50%;
|
||
text-align: center;
|
||
line-height: 32px;
|
||
font-weight: bold;
|
||
margin-left: 10px;
|
||
}
|
||
.arch-grid {
|
||
display: grid;
|
||
grid-template-columns: repeat(auto-fit, minmax(280px, 1fr));
|
||
gap: 20px;
|
||
margin: 20px 0;
|
||
}
|
||
.arch-card {
|
||
background: white;
|
||
padding: 20px;
|
||
border-radius: 10px;
|
||
box-shadow: 0 4px 12px rgba(0,0,0,0.08);
|
||
border-top: 4px solid var(--secondary);
|
||
transition: transform 0.3s;
|
||
}
|
||
.arch-card:hover { transform: translateY(-5px); }
|
||
.arch-card h4 { color: var(--primary); margin-bottom: 12px; font-size: 1.15em; }
|
||
.arch-card ul { list-style: none; padding-right: 0; }
|
||
.arch-card li { padding: 6px 0; padding-right: 20px; position: relative; font-size: 0.95em; }
|
||
.arch-card li::before { content: '▸'; position: absolute; right: 0; color: var(--accent); font-weight: bold; }
|
||
.tech-badge {
|
||
display: inline-block;
|
||
background: var(--secondary);
|
||
color: white;
|
||
padding: 4px 12px;
|
||
border-radius: 20px;
|
||
font-size: 0.85em;
|
||
margin: 3px;
|
||
}
|
||
.tech-badge.primary { background: var(--primary); }
|
||
.tech-badge.success { background: var(--success); }
|
||
.tech-badge.accent { background: var(--accent); }
|
||
pre {
|
||
background: var(--bg-code);
|
||
color: #e2e8f0;
|
||
padding: 18px;
|
||
border-radius: 8px;
|
||
overflow-x: auto;
|
||
direction: ltr;
|
||
text-align: left;
|
||
font-family: 'Consolas', 'Courier New', monospace;
|
||
font-size: 0.85em;
|
||
margin: 15px 0;
|
||
border-right: 4px solid var(--accent);
|
||
}
|
||
code {
|
||
background: #fef3c7;
|
||
color: #92400e;
|
||
padding: 2px 8px;
|
||
border-radius: 4px;
|
||
font-family: 'Consolas', monospace;
|
||
font-size: 0.9em;
|
||
direction: ltr;
|
||
display: inline-block;
|
||
}
|
||
table {
|
||
width: 100%;
|
||
border-collapse: collapse;
|
||
margin: 15px 0;
|
||
background: white;
|
||
border-radius: 8px;
|
||
overflow: hidden;
|
||
box-shadow: 0 2px 8px rgba(0,0,0,0.05);
|
||
}
|
||
th { background: var(--primary); color: white; padding: 12px; text-align: right; font-weight: bold; }
|
||
td { padding: 12px; border-bottom: 1px solid var(--border); }
|
||
tr:hover { background: var(--bg-light); }
|
||
.flow-diagram {
|
||
background: white;
|
||
padding: 25px;
|
||
border-radius: 10px;
|
||
margin: 20px 0;
|
||
text-align: center;
|
||
}
|
||
.flow-step {
|
||
display: inline-block;
|
||
background: var(--secondary);
|
||
color: white;
|
||
padding: 10px 20px;
|
||
border-radius: 8px;
|
||
margin: 5px;
|
||
font-size: 0.9em;
|
||
}
|
||
.flow-arrow { display: inline-block; color: var(--accent); font-size: 1.5em; margin: 0 8px; vertical-align: middle; }
|
||
.alert { padding: 15px 20px; border-radius: 8px; margin: 15px 0; border-right: 4px solid; }
|
||
.alert-info { background: #dbeafe; border-color: var(--secondary); color: #1e40af; }
|
||
.alert-success { background: #d1fae5; border-color: var(--success); color: #065f46; }
|
||
.alert-warning { background: #fef3c7; border-color: var(--accent); color: #92400e; }
|
||
.footer { background: var(--primary); color: white; padding: 25px; text-align: center; margin-top: 40px; }
|
||
.footer p { margin: 5px 0; }
|
||
.highlight { background: linear-gradient(120deg, #fef3c7 0%, #fef3c7 100%); padding: 2px 6px; border-radius: 4px; font-weight: bold; }
|
||
.toc { background: white; padding: 20px; border-radius: 10px; margin-bottom: 25px; border: 2px solid var(--border); }
|
||
.toc h3 { color: var(--primary); margin-bottom: 15px; }
|
||
.toc ol { padding-right: 25px; }
|
||
.toc li { padding: 6px 0; }
|
||
.toc a { color: var(--secondary); text-decoration: none; transition: color 0.2s; }
|
||
.toc a:hover { color: var(--primary); text-decoration: underline; }
|
||
.file-change { background: #f0f9ff; border-right: 4px solid var(--secondary); padding: 15px; margin: 10px 0; border-radius: 8px; }
|
||
.file-change .path { font-family: 'Consolas', monospace; color: var(--primary); font-weight: bold; direction: ltr; display: inline-block; }
|
||
.badge-new { background: var(--success); color: white; padding: 2px 8px; border-radius: 4px; font-size: 0.75em; margin-right: 8px; }
|
||
.badge-modify { background: var(--warning); color: white; padding: 2px 8px; border-radius: 4px; font-size: 0.75em; margin-right: 8px; }
|
||
</style>
|
||
</head>
|
||
<body>
|
||
<div class="container">
|
||
|
||
<div class="header">
|
||
<h1>🧠 استخراج هوشمند محتوا از فایل با مدل Qwen3.5 GGUF</h1>
|
||
<div class="subtitle">راهنمای گام به گام یکپارچهسازی مدلهای محلی GGUF در معماری xAiApi</div>
|
||
</div>
|
||
|
||
<div class="meta-bar">
|
||
<div class="meta-item">👨💻 <strong>توسعهدهنده:</strong> هادی خزاعی اصل</div>
|
||
<div class="meta-item">🏢 <strong>شرکت:</strong> فن آوران ساحر علم</div>
|
||
<div class="meta-item">📅 <strong>تاریخ تهیه مستند:</strong> چهارشنبه ۸ مهر ۱۴۰۵</div>
|
||
<div class="meta-item">📦 <strong>پروژه:</strong> xAiApi</div>
|
||
</div>
|
||
|
||
<div class="content">
|
||
|
||
<div class="toc">
|
||
<h3>📑 فهرست مطالب</h3>
|
||
<ol>
|
||
<li><a href="#overview">تحلیل رویکرد و استراتژی</a></li>
|
||
<li><a href="#step1">گام ۱: آمادهسازی مدل GGUF در Ollama</a></li>
|
||
<li><a href="#step2">گام ۲: پیکربندی مدل در appsettings.json</a></li>
|
||
<li><a href="#step3">گام ۳: ایجاد سرویس استخراج محتوا (Extraction Service)</a></li>
|
||
<li><a href="#step4">گام ۴: ایجاد Controller اختصاصی</a></li>
|
||
<li><a href="#step5">گام ۵: ثبت وابستگیها (DI)</a></li>
|
||
<li><a href="#flow">جریان کامل پردازش</a></li>
|
||
<li><a href="#tips">نکات کلیدی و مهندسی پرامپت</a></li>
|
||
</ol>
|
||
</div>
|
||
|
||
<!-- Section 1: Overview -->
|
||
<div class="section" id="overview">
|
||
<h2>🎯 گام ۱: تحلیل رویکرد و استراتژی</h2>
|
||
<p>
|
||
مدل <span class="highlight">Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF</span>
|
||
یک مدل زبانی بزرگ (LLM) در فرمت <strong>GGUF</strong> است. برای اجرای این مدل و استفاده از آن جهت استخراج محتوا،
|
||
بهترین و سازگارترین راه با معماری فعلی پروژه شما، استفاده از <strong>Ollama</strong> به عنوان موتور اجرای محلی (Local Inference Engine) است.
|
||
</p>
|
||
<div class="alert alert-info">
|
||
<strong>💡 استراتژی دو مرحلهای استخراج محتوا:</strong><br>
|
||
۱. <strong>استخراج خام (Raw Extraction):</strong> استفاده از <code>IFileContentExtractor</code> (که قبلاً طراحی کردیم) برای خواندن متن خام از PDF/DOCX/TXT.<br>
|
||
۲. <strong>پردازش هوشمند (LLM Processing):</strong> ارسال متن خام به مدل Qwen3.5 با یک <strong>System Prompt</strong> دقیق برای خلاصهسازی، استخراج موجودیتها (NER)، یا تبدیل به JSON ساختاریافته.
|
||
</div>
|
||
</div>
|
||
|
||
<!-- Section 2: Ollama Setup -->
|
||
<div class="section" id="step1">
|
||
<h2>⚙️ گام ۲: آمادهسازی مدل GGUF در Ollama</h2>
|
||
<p>از آنجا که Ollama به صورت بومی از فرمت GGUF پشتیبانی میکند، باید مدل را دانلود و در Ollama ایمپورت کنیم.</p>
|
||
|
||
<div class="step">
|
||
<span class="step-number">۱</span>
|
||
<strong>دانلود فایل GGUF:</strong>
|
||
<p>فایل <code>.gguf</code> را از لینک Hugging Face ارائه شده دانلود کرده و در مسیری مانند <code>C:\Models\qwen3.5-defiant.gguf</code> ذخیره کنید.</p>
|
||
</div>
|
||
|
||
<div class="step">
|
||
<span class="step-number">۲</span>
|
||
<strong>ایجاد فایل Modelfile:</strong>
|
||
<p>یک فایل متنی بدون پسوند به نام <code>Modelfile</code> در کنار فایل GGUF ایجاد کنید و محتوای زیر را در آن قرار دهید:</p>
|
||
<pre>FROM C:/Models/qwen3.5-defiant.gguf
|
||
|
||
# تنظیم پارامترهای بهینه برای استخراج محتوا
|
||
PARAMETER temperature 0.2
|
||
PARAMETER top_p 0.9
|
||
PARAMETER num_ctx 8192
|
||
|
||
# تنظیم پرامپت سیستمی پیشفرض برای استخراج ساختاریافته
|
||
SYSTEM """
|
||
تو یک دستیار هوشمند و دقیق برای استخراج و تحلیل محتوای اسناد هستی.
|
||
وظیفه تو خواندن متن ورودی، درک عمیق آن، و استخراج اطلاعات کلیدی به صورت ساختاریافته و دقیق است.
|
||
همیشه به زبان فارسی روان و حرفهای پاسخ بده، مگر اینکه خلاف آن درخواست شود.
|
||
"""</pre>
|
||
</div>
|
||
|
||
<div class="step">
|
||
<span class="step-number">۳</span>
|
||
<strong>ایجاد مدل در Ollama:</strong>
|
||
<p>ترمینال یا CMD را باز کرده و دستور زیر را اجرا کنید تا مدل با یک نام مستعار (Alias) کوتاه ثبت شود:</p>
|
||
<pre>ollama create qwen3.5-defiant:9b -f Modelfile</pre>
|
||
<p>سپس برای اطمینان از صحت نصب، دستور <code>ollama list</code> را اجرا کنید.</p>
|
||
</div>
|
||
</div>
|
||
|
||
<!-- Section 3: Configuration -->
|
||
<div class="section" id="step2">
|
||
<h2>📝 گام ۳: پیکربندی مدل در appsettings.json</h2>
|
||
<p>اکنون باید این مدل جدید را به لیست مدلهای مجاز در پیکربندی پروژه اضافه کنیم تا سرویسها بتوانند از آن استفاده کنند.</p>
|
||
|
||
<div class="file-change">
|
||
<span class="badge-modify">MODIFY</span>
|
||
<span class="path">xAiApi/appsettings.json</span>
|
||
</div>
|
||
|
||
<pre>
|
||
"AiApiConfiguration": {
|
||
"Models": [
|
||
{
|
||
"Name": "Gemma",
|
||
"Url": "http://localhost:11434",
|
||
"LLM": "gemma:2b",
|
||
"Provider": "Ollama"
|
||
},
|
||
{
|
||
"Name": "QwenDefiant",
|
||
"Url": "http://localhost:11434",
|
||
"LLM": "qwen3.5-defiant:9b",
|
||
"Provider": "Ollama"
|
||
}
|
||
],
|
||
"Prompts": [
|
||
{
|
||
"Name": "Introduction",
|
||
"Template": "تو یک دستیار هوشمند مفید هستی."
|
||
},
|
||
{
|
||
"Name": "DocumentExtraction",
|
||
"Template": "متن زیر از یک فایل استخراج شده است. لطفاً آن را تحلیل کن و خروجی را دقیقاً در قالب JSON با کلیدهای زیر برگردان: {\"summary\": \"خلاصه ۳ خطی\", \"key_points\": [\"نکته ۱\", \"نکته ۲\"], \"entities\": {\"نام_اشخاص\": [], \"تاریخ_ها\": []}}. متن: {0}"
|
||
}
|
||
]
|
||
}
|
||
</pre>
|
||
<div class="alert alert-success">
|
||
<strong>✅ نکته:</strong> استفاده از نام مستعار <code>qwen3.5-defiant:9b</code> در فیلد <code>LLM</code> باعث میشود کد شما تمیز و خوانا بماند.
|
||
</div>
|
||
</div>
|
||
|
||
<!-- Section 4: Extraction Service -->
|
||
<div class="section" id="step3">
|
||
<h2>🔧 گام ۴: ایجاد سرویس استخراج محتوا (Extraction Service)</h2>
|
||
<p>ما یک سرویس اختصاصی ایجاد میکنیم که ترکیبی از <code>IFileContentExtractor</code> (برای خواندن فایل) و <code>IChatClient</code> (برای پردازش با Qwen) باشد.</p>
|
||
|
||
<div class="file-change">
|
||
<span class="badge-new">NEW</span>
|
||
<span class="path">xAiApi/Interfaces/IDocumentExtractionService.cs</span>
|
||
</div>
|
||
|
||
<pre>using System.Threading;
|
||
using System.Threading.Tasks;
|
||
using Microsoft.AspNetCore.Http;
|
||
|
||
namespace xAiApi.Interfaces
|
||
{
|
||
public interface IDocumentExtractionService
|
||
{
|
||
/// <summary>
|
||
/// استخراج و تحلیل هوشمند محتوای یک فایل
|
||
/// </summary>
|
||
Task<string> ExtractAndAnalyzeAsync(
|
||
IFormFile file,
|
||
string extractionPromptTemplate,
|
||
CancellationToken cancellationToken = default
|
||
);
|
||
}
|
||
}</pre>
|
||
|
||
<div class="file-change">
|
||
<span class="badge-new">NEW</span>
|
||
<span class="path">xAiApi/Providers/XDocumentExtractionService.cs</span>
|
||
</div>
|
||
|
||
<pre>using System;
|
||
using System.IO;
|
||
using System.Threading;
|
||
using System.Threading.Tasks;
|
||
using Microsoft.AspNetCore.Http;
|
||
using Microsoft.Extensions.AI;
|
||
using Microsoft.Extensions.Logging;
|
||
using xAiApi.Configurations;
|
||
using xAiApi.Constants;
|
||
using xAiApi.Extensions;
|
||
using xAiApi.Interfaces;
|
||
using xAiModels.Constants;
|
||
using xAiModels.Extensions;
|
||
using xAiService.Interfaces;
|
||
using xExceptions.Constants;
|
||
|
||
namespace xAiApi.Providers
|
||
{
|
||
public class XDocumentExtractionService : IDocumentExtractionService
|
||
{
|
||
private readonly IFileContentExtractor _fileExtractor;
|
||
private readonly XAiApiConfiguration _configuration;
|
||
private readonly ILogger<XDocumentExtractionService> _logger;
|
||
|
||
public XDocumentExtractionService(
|
||
IFileContentExtractor fileExtractor,
|
||
XAiApiConfiguration configuration,
|
||
ILogger<XDocumentExtractionService> logger)
|
||
{
|
||
_fileExtractor = fileExtractor;
|
||
_configuration = configuration;
|
||
_logger = logger;
|
||
}
|
||
|
||
public async Task<string> ExtractAndAnalyzeAsync(
|
||
IFormFile file,
|
||
string extractionPromptTemplate,
|
||
CancellationToken cancellationToken = default)
|
||
{
|
||
// ۱. اعتبارسنجی فایل
|
||
if (file == null || file.Length == 0)
|
||
{
|
||
XException.InvalidArgs.Throw("فایل نامعتبر است.");
|
||
}
|
||
|
||
// ۲. استخراج متن خام از فایل
|
||
string rawText;
|
||
using (var stream = file.OpenReadStream())
|
||
{
|
||
if (_fileExtractor.CanExtract(file.ContentType))
|
||
{
|
||
rawText = await _fileExtractor.ExtractAsync(stream, file.ContentType, cancellationToken);
|
||
}
|
||
else
|
||
{
|
||
XException.InvalidData.Throw($"فرمت فایل {file.ContentType} پشتیبانی نمیشود.");
|
||
}
|
||
}
|
||
|
||
if (string.IsNullOrWhiteSpace(rawText))
|
||
{
|
||
XException.InvalidData.Throw("محتوای استخراج شده از فایل خالی است.");
|
||
}
|
||
|
||
// ۳. آمادهسازی پرامپت نهایی
|
||
var finalPrompt = string.Format(extractionPromptTemplate, rawText);
|
||
|
||
// ۴. دریافت کلاینت مدل QwenDefiant از پیکربندی
|
||
var modelDescriptor = _configuration.GetModel("QwenDefiant");
|
||
if (!modelDescriptor.IsValid())
|
||
{
|
||
XException.InvalidConfiguration.Throw("مدل QwenDefiant در پیکربندی یافت نشد.");
|
||
}
|
||
|
||
// ۵. ساخت ChatClient (با استفاده از منطق موجود در XAIServiceBase یا مستقیم)
|
||
using var client = CreateChatClient(modelDescriptor);
|
||
|
||
// ۶. ارسال درخواست به مدل
|
||
var history = new[]
|
||
{
|
||
new ChatMessage(ChatRole.System, "تو یک متخصص استخراج داده از اسناد هستی. فقط خروجی درخواست شده را تولید کن."),
|
||
new ChatMessage(ChatRole.User, finalPrompt)
|
||
};
|
||
|
||
var response = await client.GetResponseAsync(history, cancellationToken: cancellationToken);
|
||
|
||
if (!response.IsValid())
|
||
{
|
||
XException.ActionFailed.Throw("مدل هوش مصنوعی پاسخی معتبر تولید نکرد.");
|
||
}
|
||
|
||
return response.Text;
|
||
}
|
||
|
||
private IChatClient CreateChatClient(XAiModels.Models.XAiModelDescriptor descriptor)
|
||
{
|
||
// بازنویسی منطق ساخت کلاینت Ollama بر اساس معماری پروژه
|
||
var httpClient = new HttpClient { BaseAddress = new Uri(descriptor.Url), Timeout = Timeout.InfiniteTimeSpan };
|
||
var ollamaClient = new OllamaSharp.OllamaApiClient(httpClient, descriptor.LLM);
|
||
|
||
return new ChatClientBuilder(ollamaClient)
|
||
.UseFunctionInvocation()
|
||
.Build();
|
||
}
|
||
}
|
||
}</pre>
|
||
</div>
|
||
|
||
<!-- Section 5: Controller -->
|
||
<div class="section" id="step4">
|
||
<h2>🎮 گام ۵: ایجاد Controller اختصاصی</h2>
|
||
<p>یک endpoint جدید برای دریافت فایل و بازگرداندن محتوای تحلیلشده ایجاد میکنیم.</p>
|
||
|
||
<div class="file-change">
|
||
<span class="badge-new">NEW</span>
|
||
<span class="path">xAiApi/Controllers/DocumentExtractionController.cs</span>
|
||
</div>
|
||
|
||
<pre>using System;
|
||
using System.Threading;
|
||
using System.Threading.Tasks;
|
||
using Microsoft.AspNetCore.Mvc;
|
||
using Microsoft.AspNetCore.Authorization;
|
||
using Microsoft.Extensions.Logging;
|
||
using xCommons.Configurations;
|
||
using xCommons.Providers;
|
||
using xIdentityService.Interfaces;
|
||
using xAiApi.Interfaces;
|
||
using xAiApi.Extensions;
|
||
|
||
namespace xAiApi.Controllers
|
||
{
|
||
[Authorize]
|
||
[Route("api/[controller]")]
|
||
public class DocumentExtractionController : XBaseIdentityApiV1Controller
|
||
{
|
||
private readonly IDocumentExtractionService _extractionService;
|
||
private readonly XAiApiConfiguration _configuration;
|
||
|
||
public DocumentExtractionController(
|
||
ILogger<DocumentExtractionController> logger,
|
||
XAppConfiguration appConfiguration,
|
||
XValidationProvider validationProvider,
|
||
IXIdentityProvider identityProvider,
|
||
IDocumentExtractionService extractionService,
|
||
XAiApiConfiguration configuration)
|
||
: base(logger, appConfiguration, validationProvider, identityProvider)
|
||
{
|
||
_extractionService = extractionService;
|
||
_configuration = configuration;
|
||
}
|
||
|
||
/// <summary>
|
||
/// آپلود فایل و استخراج هوشمند محتوا با مدل Qwen3.5
|
||
/// </summary>
|
||
[HttpPost("Extract")]
|
||
[Consumes("multipart/form-data")]
|
||
public async Task<ActionResult<object>> Extract(
|
||
[FromForm] IFormFile file,
|
||
[FromForm] string promptName = "DocumentExtraction",
|
||
CancellationToken cancellationToken = default)
|
||
{
|
||
try
|
||
{
|
||
ValidationProvider.NotNull(file, nameof(file));
|
||
|
||
// دریافت الگوی پرامپت از پیکربندی
|
||
var promptTemplate = _configuration.GetPrompt(promptName);
|
||
if (string.IsNullOrWhiteSpace(promptTemplate))
|
||
{
|
||
promptTemplate = "متن زیر را تحلیل و خلاصه کن: {0}";
|
||
}
|
||
|
||
// فراخوانی سرویس استخراج
|
||
var result = await _extractionService.ExtractAndAnalyzeAsync(
|
||
file: file,
|
||
extractionPromptTemplate: promptTemplate,
|
||
cancellationToken: cancellationToken
|
||
);
|
||
|
||
return Ok(new { success = true, data = result });
|
||
}
|
||
catch (Exception ex)
|
||
{
|
||
return GetExceptionActionResult(ex);
|
||
}
|
||
}
|
||
}
|
||
}</pre>
|
||
</div>
|
||
|
||
<!-- Section 6: DI -->
|
||
<div class="section" id="step5">
|
||
<h2>🔗 گام ۶: ثبت وابستگیها (Dependency Injection)</h2>
|
||
<p>سرویس جدید را در متد <code>ConfigureServices</code> فایل <code>Startup.cs</code> ثبت کنید.</p>
|
||
|
||
<div class="file-change">
|
||
<span class="badge-modify">MODIFY</span>
|
||
<span class="path">xAiApi/Startup.cs</span>
|
||
</div>
|
||
|
||
<pre>public void ConfigureServices(IServiceCollection services)
|
||
{
|
||
// ... (ثبتهای قبلی)
|
||
|
||
// ثبت Extractor های فایل (اگر قبلاً ثبت نشدهاند)
|
||
services.AddSingleton<IFileContentExtractor, PlainTextContentExtractor>();
|
||
services.AddSingleton<IFileContentExtractor, PdfContentExtractor>();
|
||
services.AddSingleton<IFileContentExtractor, CompositeFileContentExtractor>();
|
||
|
||
// ✅ ثبت سرویس جدید استخراج محتوا
|
||
services.AddScoped<IDocumentExtractionService, XDocumentExtractionService>();
|
||
|
||
// ... (بقیه کدها)
|
||
}</pre>
|
||
</div>
|
||
|
||
<!-- Section 7: Flow -->
|
||
<div class="section" id="flow">
|
||
<h2>🔄 گام ۷: جریان کامل پردازش</h2>
|
||
<div class="flow-diagram">
|
||
<span class="flow-step">📤 کلاینت: آپلود فایل</span>
|
||
<span class="flow-arrow">→</span>
|
||
<span class="flow-step">🎮 DocumentExtractionController</span>
|
||
<span class="flow-arrow">→</span>
|
||
<span class="flow-step">📄 IFileContentExtractor (استخراج متن خام)</span>
|
||
<span class="flow-arrow">→</span>
|
||
<span class="flow-step">🧠 Qwen3.5 GGUF (تحلیل و ساختارسازی)</span>
|
||
<span class="flow-arrow">→</span>
|
||
<span class="flow-step">✅ بازگرداندن JSON/متن تحلیلشده</span>
|
||
</div>
|
||
|
||
<h3>نمونه درخواست (cURL):</h3>
|
||
<pre>curl -X POST "http://localhost:5000/api/DocumentExtraction/Extract" \
|
||
-H "Authorization: Bearer YOUR_TOKEN" \
|
||
-F "file=@/path/to/document.pdf" \
|
||
-F "promptName=DocumentExtraction"</pre>
|
||
</div>
|
||
|
||
<!-- Section 8: Tips -->
|
||
<div class="section" id="tips">
|
||
<h2>💡 گام ۸: نکات کلیدی و مهندسی پرامپت برای مدلهای GGUF</h2>
|
||
|
||
<div class="arch-grid">
|
||
<div class="arch-card">
|
||
<h4>🎯 پرامپتنویسی دقیق</h4>
|
||
<p>مدلهای GGUF محلی به دستورالعملهای شفاف بسیار خوب پاسخ میدهند. در <code>appsettings.json</code> حتماً قالب خروجی (مثلاً JSON) را به صراحت مشخص کنید.</p>
|
||
</div>
|
||
<div class="arch-card">
|
||
<h4>⚡ مدیریت Context Window</h4>
|
||
<p>در <code>Modelfile</code> مقدار <code>num_ctx</code> را بر اساس حجم فایلهای شما تنظیم کنید (مثلاً 8192 یا 16384). اگر فایل بزرگ است، آن را به قطعات (Chunks) تقسیم کنید.</p>
|
||
</div>
|
||
<div class="arch-card">
|
||
<h4>🛡️ مدیریت خطا</h4>
|
||
<p>همیشه احتمال خطای OOM (کمبود حافظه RAM/VRAM) در مدلهای 9B را در نظر بگیرید. لاگهای Ollama را برای پایش مصرف حافظه بررسی کنید.</p>
|
||
</div>
|
||
<div class="arch-card">
|
||
<h4>🖼️ پشتیبانی از تصویر (Vision)</h4>
|
||
<p>اگر این نسخه خاص از Qwen از ورودی تصویر پشتیبانی کند، میتوانید در <code>XDocumentExtractionService</code> به جای متن خام، فایل تصویر را به <code>DataContent</code> تبدیل و ارسال کنید.</p>
|
||
</div>
|
||
</div>
|
||
|
||
<div class="alert alert-success">
|
||
<strong>✅ جمعبندی:</strong> با این طراحی، شما بدون تغییر در هسته اصلی <code>XAIServiceBase</code>، یک ماژول کاملاً ایزوله و قدرتمند برای استخراج محتوا با مدلهای محلی GGUF ایجاد کردهاید که کاملاً با معماری ماژولار شرکت <strong>فن آوران ساحر علم</strong> همخوانی دارد.
|
||
</div>
|
||
</div>
|
||
|
||
</div>
|
||
|
||
<div class="footer">
|
||
<p><strong>👨💻 توسعهدهنده:</strong> هادی خزاعی اصل</p>
|
||
<p><strong>🏢 شرکت:</strong> فن آوران ساحر علم</p>
|
||
<p><strong>📅 تاریخ تهیه مستند:</strong> چهارشنبه ۸ مهر ۱۴۰۵</p>
|
||
<p style="margin-top: 15px; opacity: 0.8; font-size: 0.9em;">
|
||
🧠 طراحی سیستم استخراج محتوا با Qwen3.5 GGUF - تمامی حقوق محفوظ است
|
||
</p>
|
||
</div>
|
||
|
||
</div>
|
||
</body>
|
||
</html> |