Web playgrounds are great for demos. You paste a block of text, watch the model generate a neat summary, and close the tab. But that is not engineering. Production work means APIs, error handling, and code that runs while you sleep. If you need to churn through meeting transcripts, support tickets, or research papers on a schedule, you need a pipeline.
This guide walks through building exactly that: a lightweight, automated document summarization script using Python, the AWS SDK for Python (boto3), and Amazon Bedrock. We will use Anthropic’s Claude 3 Haiku, a model that hits the sweet spot of speed and cost for text summarization tasks.
Why Bedrock and Claude 3 Haiku?
Amazon Bedrock is a managed service that exposes foundation models through a single set of AWS APIs. Instead of piecing together external endpoints and wrestling with separate billing and security models, you call an AWS endpoint with standard IAM controls. Your data stays within your AWS environment.
Claude 3 Haiku is the leanest model in Anthropic’s Claude 3 family. It is built for responsiveness and low cost, which makes it ideal for high-volume summarization where you want predictable output without paying for the horsepower of larger models on simple reading tasks.
What You Need
Before writing any code, make sure you have the following ready:
- An active AWS account.
- Python 3.9 or higher installed locally.
- The AWS CLI configured with credentials that have permission to invoke Bedrock models. If you have not run
aws configureyet, do that now. If you hit permission errors later, you will likely need to attach the appropriate Bedrock invocation permissions to your IAM user or role. - Model access enabled specifically for Anthropic Claude 3 Haiku inside the AWS Bedrock console. AWS requires you to explicitly opt in for each model provider before you can call it.
Step 1: Enable Model Access
Bedrock does not let you call models out of the box. You must flip the switch in the console first.
- Log into the AWS Management Console.
- Use the search bar to find Amazon Bedrock.
- In the left navigation panel, select Model access.
- Click Modify model access.
- Tick the box for Anthropic (Claude 3 Haiku) and submit your request.
Once the status flips to "Access granted," you are clear to call the model from code.
Step 2: Set Up Your Environment
A clean Python environment keeps dependencies isolated and reproducible. Open your terminal and run these commands:
mkdir bedrock-summarizer && cd bedrock-summarizer
python3 -m venv venv
source venv/bin/activate
pip install boto3
Windows users should replace the activation command with venv\Scripts\activate. After pip install boto3 finishes, you have everything you need to talk to AWS APIs.
Step 3: Write the Script
Create a file named summarize.py. The goal is to read a document from disk, hand it to the Bedrock Converse API, and print a concise summary.
Below is a complete, working implementation. We use the Converse API because it abstracts away the raw JSON formatting that different model providers expect. You simply pass a list of messages and inference settings.
import boto3
def summarize_document(text: str) -> str:
client = boto3.client("bedrock-runtime")
model_id = "anthropic.claude-3-haiku-20240307-v1:0"
messages = [
{
"role": "user",
"content": [
{
"text": (
"Provide a concise summary of the following document. "
"Focus on the main points and avoid unnecessary detail:\n\n"
f"{text}"
)
}
]
}
]
response = client.converse(
modelId=model_id,
messages=messages,
inferenceConfig={
"temperature": 0.3,
"maxTokens": 512
}
)
summary = response["output"]["message"]["content"][0]["text"]
return summary.strip()
if __name__ == "__main__":
with open("document.txt", "r", encoding="utf-8") as f:
document_text = f.read()
result = summarize_document(document_text)
print("\n--- Summary ---\n")
print(result)
A few practical details worth highlighting here:
- boto3.client("bedrock-runtime") targets the runtime endpoint that handles inference. Make sure your AWS region in
~/.aws/configsupports Bedrock and that you enabled Haiku in that same region. - Model ID
anthropic.claude-3-haiku-20240307-v1:0is the exact identifier Bedrock expects. Copy it precisely. - Temperature set to 0.3 keeps the output grounded. For summarization, you want consistency and fidelity to the source text, not creative embellishment. If you raise the temperature toward 1.0, the model starts taking liberties with phrasing and occasionally invents details.
- The prompt itself is specific. Instead of throwing raw text at the model with a vague "summarize this," we explicitly ask for main points and instruct it to skip fluff. That kind of clarity separates unusable output from something you can actually ship.
Place any text file you want to summarize in the same directory and name it document.txt.
Step 4: Run It
With your virtual environment active, execute:
python summarize.py
If your credentials and model access are correct, you should see a tidy summary printed to your terminal within a few seconds. If you get an access error, double-check your IAM permissions and confirm you enabled Claude 3 Haiku in the console.
Pushing Beyond the Script
এই পাইপলাইনটি উদ্দেশ্যমূলকভাবে সহজ রাখা হয়েছে, তবে এটি প্রকৃত অটোমেশনের ভিত্তি। অতিরিক্ত জটিলতা না বাড়িয়ে আপনি কীভাবে এটিকে আরও উন্নত করতে পারেন তা নিচে দেওয়া হলো।
Batch processing. একটি একক ফাইল পড়ার পরিবর্তে একটি ডিরেক্টরির ওপর লুপ ব্যবহার করুন। একটি ইনপুট ফোল্ডারে পঞ্চাশটি PDF বা টেক্সট ফাইল রাখুন, সেগুলোর মধ্য দিয়ে ইটারেট (iterate) করুন এবং সামারিগুলো একটি আউটপুট ফোল্ডারে লিখুন। আপনি যদি সরাসরি PDF ইনজেস্ট করতে চান, তবে Bedrock-এ পাঠানোর আগে র (raw) টেক্সট এক্সট্র্যাক্ট করার জন্য PyPDF2 বা pdfplumber-এর মতো লাইব্রেরি দিয়ে একটি প্রি-প্রসেসিং ধাপের প্রয়োজন হবে।
Chunking strategy. খুব দীর্ঘ ডকুমেন্ট মডেলের কনটেক্সট লিমিট (context limit) অতিক্রম করতে পারে। এমন ক্ষেত্রে, টেক্সটটিকে অনুচ্ছেদ বা সেকশন অনুযায়ী যৌক্তিক চাঙ্কে (chunks) বিভক্ত করুন, প্রতিটি চাঙ্ক আলাদাভাবে সামারি করুন এবং তারপর চূড়ান্ত সংশ্লেষণের (synthesis) জন্য মধ্যবর্তী সামারিগুলো পুনরায় মডেলের কাছে পাঠান। এই দ্বি-স্তরীয় পদ্ধতিটি পুরো ডকুমেন্টের বিষয়বস্তু বজায় রেখে আপনাকে টোকেন লিমিটের মধ্যে রাখতে সাহায্য করবে।
Error handling. প্রোডাকশন কোডে বিশেষভাবে boto3.exceptions.ClientError হ্যান্ডেল করা উচিত। আপনি যদি খুব দ্রুত বা অতিরিক্ত মাত্রায় API কল করেন, তবে AWS আপনার রিকোয়েস্ট থ্রটল (throttle) করতে পারে। আপনার converse কলটিকে 'exponential backoff' সহ একটি রিট্রাই লুপে রাখুন, অথবা রেট লিমিটগুলো সুন্দরভাবে সামলানোর জন্য tenacity-এর মতো লাইব্রেরি ব্যবহার করুন।
Prompt engineering. একটি সাধারণ সামারি এবং একটি কার্যকর সামারির মধ্যে পার্থক্যটি প্রায়ই প্রম্পটের ওপর নির্ভর করে। যদি আপনি দ্রুত পড়ার উপযোগী (scanability) কিছু চান, তবে বুলেট পয়েন্টের জন্য অনুরোধ করুন। যদি শ্রোতা উচ্চপদস্থ কর্মকর্তা হন, তবে এক অনুচ্ছেদের একটি এক্সিকিউটিভ সামারি চান। আপনি এমনকি ফরম্যাটিংয়ের সীমাবদ্ধতাও দিতে পারেন, যেমন "সামারিটি তিন বাক্যের মধ্যে সীমাবদ্ধ রাখুন" অথবা "topic, key_points, এবং action_items কী (key) সহ JSON হিসেবে আউটপুট প্রদান করুন।"
আসল শিক্ষা
একটি চ্যাট প্লেগ্রাউন্ড থেকে একটি কার্যকরী স্ক্রিপ্টে উত্তরণ হলো সেই সন্ধিক্ষণ যেখানে AI একটি ইনফ্রাস্ট্রাকচারে পরিণত হয়। একবার এই পাইপলাইনটি লোকালি চলতে শুরু করলে, আপনি এটিকে S3 আপলোডের মাধ্যমে ট্রিগার হওয়া একটি AWS Lambda ফাংশনে নিয়ে যেতে পারেন, ECS Fargate-এ শিডিউল করতে পারেন, অথবা বিদ্যমান কোনো ডেটা ওয়ার্কফ্লোতে যুক্ত করতে পারেন। API কল করা সহজ অংশ। আসল ইঞ্জিনিয়ারিং ভ্যালু আসে সেই কলটিকে এমন লজিকের মাধ্যমে আবৃত করার মাধ্যমে যা ফাইল, এরর এবং ফরম্যাটিং সামলাতে পারে, যাতে আপনাকে আর কখনোই ব্রাউজারে টেক্সট কপি এবং পেস্ট করতে না হয়।
এই সেটআপ সম্পর্কে আরও বিস্তারিত এবং বিভিন্ন সংস্করণ জানতে Dev.to-তে মূল ওয়াকথ্রু (walkthrough) দেখুন। আপনি যদি নির্মাতাদের (builders) একটি কমিউনিটির সাথে AWS আর্কিটেকচার, LLM পাইপলাইন বা প্রম্পট ইঞ্জিনিয়ারিং নিয়ে আলোচনা করতে চান, তবে [GyaanSetu AI on Telegram](https://t.me/GyaanSet
