Switch language한국어
Back to the list

Case Study on Optimizing AWS g6e-Based LLM Inference Batch Workloads of Neosapiens

TL;DR AI

Key summary

2 min read
  1. Neosapience operates an AI actor service called Typecast based on AI voice synthesis and language intelligence technology.

  2. Since its establishment in 2017, it has researched deep learning-based emotional expression and multilingual TTS technologies.

  3. This post analyzes the optimization of lightweight LLM inference on Amazon EC2 instances, focusing on throughput and latency.

  4. The combination of g6e (L40S) and INT8 was found to provide the best balance for operational conditions.

  5. Neosapience aims to refine its optimal instance-precision combination in light of evolving AI product trends.

Read the original