Srivatsa Koustub

Backend, AI & Machine Learning Engineer

ƒ +91 (863) 959-6953 # kousthub.srivatsa@gmail.com ï linkedin.com/in/srivatsa-koustub Ðleetcode.com/u/user7293pG

About

Results-driven Backend & Machine Learning Engineer with 1.10+ years of hands-on experience in Python, Flask, Django, and REST API development. Proven expertise in building production-grade backend systems, implementing DevOps practices across full development cycles, and building and integrating AI/ML capabilities into production backend systems. Exceptional written and verbal communication skills with confident presentation abilities. Experienced in requirement gathering, technical support, and collaboration with cross-functional teams to drive measurable business results through AI-powered solutions and scalable backend architectures.

Education

  • Jain University August 2020 – June 2024
  • Bachelor of Technology in Computer Science and Engineering - Artificial Intelligence CGPA: 8.15/10.0
  • Core Technical Skills
  • Programming & Coding : Python (PEP8)
  • Backend Frameworks : Django, Flask, REST APIs, Microservices Architecture
  • Databases & Big Data : MySQL, PostgreSQL, MongoDB, SQLite, Apache Hadoop, Apache Spark, PySpark, HDFS,
  • Amazon S3
  • Cloud & DevOps : Microsoft Azure, Docker, Jenkins, CI/CD Pipelines, Git, GitHub, YARN, MapReduce
  • Analytics & AI : Machine Learning, Deep Learning, NLP, RAG, Big Data Processing, Business Analytics, TensorFlow,
  • Scikit-Learn, XGBoost, LangChain, Hugging Face
  • Enterprise Solutions : API Design, Web Architecture, Scalability, JWT Authentication, Production Systems Management
  • Development Practices : Agile, SDLC, Version Control, OOP, Code Auditing, Technical Feasibility Assessment

Experience

Software EngineerSeptember 2024 – Present
  • Cappitall Want Network (FinTech), Bangalore
  • Analytics Solutions & Business Strategy: Developed comprehensive GST data analytics and business reporting solutions using Python and Pandas to process 1M+ enterprise records across GSTR2A, GSTR2B, and GSTR1, implementing reconciliation logic and business strategy optimization that reduced manual auditing time by 70% and enhanced regulatory compliance through automated quality checks
  • Backend Enterprise Solutions: Designed and deployed scalable backend systems using Flask and Django frameworks, implementing REST APIs with Swagger documentation for GST compliance workflows including TDS processing (24Q and 26Q filing), MSME reporting, and IMS integration enabling direct push and reset of invoice records to the government GST portal
  • Bulk Data Processing & Performance Optimization: Engineered high-performance upload processing pipelines for sales and purchases modules handling lakhs of records, optimizing batch insert and delete operations to significantly reduce processing time and improve system throughput for large-scale enterprise data ingestion
  • Third-Party API & Government Portal Integration: Integrated third-party financial services including Vayana and Perfios APIs for real-time data enrichment, alongside government portal APIs for GSTIN/PAN validation and automated GST return synchronization, eliminating manual data entry and ensuring 99.9% data accuracy
  • Cloud DevOps & Infrastructure Management: Implemented comprehensive DevOps practices across Azure cloud ecosystem, utilizing Docker containerization, Jenkins, and CI/CD pipelines to ensure high availability and seamless automated deployments for enterprise FinTech solutions
  • AI-Powered Document Extraction System: Engineered a Gemini-based intelligent document parsing pipeline to replace brittle, layout-dependent PDF extraction (Tabula), enabling reliable field extraction from government TDS challan documents despite frequent changes in PDF structure; integrated Azure Blob Storage for document retrieval and designed dynamic JSON schemas to extract structured data with configurable field-level instructions, significantly reducing extraction failures and manual intervention
  • Tech Stack: Python, Django, Flask, MySQL, Pandas, SQLAlchemy, Swagger, Azure Queue, Docker, Jenkins, CI/CD,
  • REST APIs, JWT Authentication

Projects

  • Customer Churn Prediction using Big Data & Machine Learning | Hadoop, PySpark, Scikit-Learn, TensorFlow
  • Built end-to-end ML pipeline processing 500MB+ customer data using Apache Hadoop (HDFS) and PySpark for distributed computing, handling 1M+ records efficiently
  • Implemented data preprocessing and feature engineering on distributed datasets using PySpark transformations including window functions, aggregations, and SQL operations
  • Developed classification models achieving 87% accuracy using ensemble methods (Random Forest, Gradient Boosting) and deep learning (TensorFlow Neural Network)
  • Engineered scalable data pipeline with HDFS storage, PySpark ETL processes, and automated model training workflows reducing processing time by 60%
  • Deployed Dockerized Hadoop cluster with 3-node configuration (NameNode, DataNodes, YARN) for distributed processing and fault tolerance
  • Performed comprehensive model evaluation using precision, recall, F1-score, and ROC-AUC metrics with cross-validation for robust performance assessment
  • MultiPDF Chat App with RAG | Python, LangChain, GPT-3.5/GPT-4, Hugging Face, Vector DB
  • Developed an intelligent document QA system using Retrieval-Augmented Generation (RAG) architecture with
  • LangChain framework and OpenAI GPT models
  • Implemented document parsing and text extraction pipelines for processing multiple PDF files with chunking strategies for optimal context retrieval
  • Integrated vector embeddings using Hugging Face models for semantic search, enabling accurate document retrieval based on query similarity
  • Leveraged open-source LLMs including GPT-3.5 and GPT-4 to create context-aware conversational AI with document grounding capabilities
  • Credit Card Fraud Detection | Python, Scikit-Learn, Pandas, Imbalanced-Learn
  • Built end-to-end machine learning pipeline for detecting fraudulent credit card transactions using ensemble methods and anomaly detection algorithms
  • Addressed class imbalance using SMOTE and undersampling techniques, improving model recall for minority class by
  • 35%
  • Optimized model performance focusing on Area Under the Precision-Recall Curve (AUPRC) metric, achieving 0.89
  • AUPRC score
  • Evaluated multiple algorithms including Random Forest, XGBoost, and Logistic Regression with cross-validation and hyperparameter tuning
  • Professional Development
  • Technical Mentor | Coding Bootcamp
  • Mentored junior developers in Python coding best practices, backend development fundamentals, and problem-solving techniques, conducting code reviews and maintaining knowledge base documentation to accelerate learner progress
  • AWS Cloud Security Certification | Amazon Web Services
  • Completed comprehensive cybersecurity bootcamp by Amazon Web Services, specializing in secure coding practices, cloud security architecture, and enterprise-grade security implementations for production systems
  • Stakeholder Communication & Technical Leadership | Cappitall Want Network
  • Led collaboration with business stakeholders through confident written and verbal communication, conducting requirement gathering sessions, technical feasibility assessments, and delivering presentations showcasing AI-powered
  • FinTech solutions to cross-functional teams

Certifications

  • MuleSoft Certified Developer - Level 1 | MuleSoft, Salesforce
  • Completed industry-recognized certification validating expertise in API-led connectivity, integration design, and building scalable Mule applications for enterprise system integration
  • Types of Cyber Security | Great Learning
  • Completed professional certification covering core cybersecurity domains including network security, application security, secure coding practices, and threat mitigation strategies