Model Accuracy
98.4%
Cloud Nodes
AWS / Online
Juan Sebastián Peña
Specialized in data architecture, distributed ETL pipelines (AWS / PySpark / Airflow), and production-grade AI / Computer Vision integration.
Project Portfolio & Case Studies
Data Science and Artificial Intelligence (AI) Specialist with over 2 years of experience developing production-ready software solutions in ecosystems such as AWS and Azure. Expert in machine learning models, deep learning, computer vision algorithms, natural language processing (NLP), RAG systems, and LLM models, with experience developing ETL algorithms using orchestrators such as Apache Airflow and applying data processing tools such as Apache Spark (PySpark). Contact me for more information about my services.
Medical Implants Traceability System (DMI-Database-AI)
Digitization and migration of legacy paper and Excel-based infrastructure to an SQL database architecture for the traceability of implantable medical devices (IMDs).
🎯 Results and impact:
- 100% digitization of implant cards
- 96% software adoption rate among medical staff
- Data Validation at the Time of Entry
- Cloud-based architecture for scalability and security
Real-Time Skeleton Tracking & Pose Estimation
Computer-vision-driven upper-limb rehabilitation monitoring using gRPC microservices, ViTPose+, and automated PDF clinical reporting.
🎯 Results and impact:
- 17 Keypoints (COCO-17 Scale)
- Port 50051 (gRPC Stream)
- 0 - 90° Range Scale
- 100% Automated CI/CD
AWS Medallion Data Lake & PySpark Pipeline
Distributed data engineering pipeline with Medallion architecture (Bronze/Silver/Gold), orchestrated with Airflow and queries in Athena.
🎯 Results and impact:
- 45% Reduction in Query Latency in Athena
- Handling Small Files: Problem with `coalesce()`
- Broadcast Joins for Shuffle Elimination
- Partitioning and Bucketing for Optimized Query Performance
Biomedical Auditory and AI RAG system
Chatbot using a RAG-type system for the harmonization of Colombian and international regulations on medical devices and biomedical engineering. Implemented using modern tools and can be applied to any kind of science.
🎯 Results and impact:
- Use of vector databases for information efficiency
- 35% faster and more efficient than other similar systems.
- Deployment on AWS servers for front-end and back-end
- Use of Groq LLM for orchestration and LangChain for pipeline logic
AudioMAE Masked Autoencoders that listen - Implementation
This project implements the inference pipeline for AudioMAE (Masked Autoencoders that Listen), a state-of-the-art self-supervised model architecture presented at NeurIPS 2022 by Meta AI.
🎯 Results and impact:
- Use of model pretrained and fine-tuned for audio classification tasks
- Easy to replicate and adapt to other audio classification tasks
- Accuracy of 92% in the classification of environmental sounds and bioacoustic signals
- AudioMAE architecture with spectrogram embeddings and random masking ratios
Classification of Breast Tissue using Radiomics Algorithms
This is a radiomics-based project focused on extracting quantitative imaging biomarkers from mammography scans to support supervised classification tasks.
🎯 Results and impact:
- Training phase adquired 89.1% accuracy within test data
- KNN Neighbors model used for training purposes
- Made it possible to distinguish between three types of tissue (normal, benign, and malignant)
Cephalometric Landmark Detection - Ricketts Line Autodetection
Automatic detection of nasal and soft pogonion points in cephalometric images and visualization of the Ricketts aesthetic line using image processing and optimization techniques.
🎯 Results and impact:
- Identification of anatomical points of interest within a margin of error of 3 mm
- Automatic method with low RAM usage and no AI model integrated
- Dental practice using traditional computer vision
Professional & Research Experience
Select a node from the terminal menu to inspect key deliverables, tech stacks, and impact metrics.
AI Specialist & Data Engineer
Subocol S.A.
Developing services for enterprise software based on data engineering and artificial intelligence techniques (Machine and Deep Learning) to create distinctive service offerings that substantially improve key aspects—such as the user experience in production environments, internal process optimization, and applicable automation.
Key Deliverables & Contributions
- ›Manage projects for the development of data- and machine learning-based services across all stages, in accordance with requirements.
- ›Design and execute experimentation processes to validate proposed hypotheses.
- ›Develop and deploy production-grade LLM/RAG pipelines, PySpark data lakehouses, and containerized microservices in AWS.
- ›Perform maintenance, monitoring, and continuous improvement of already developed solutions, taking into account business feedback and data fluctuations.
Technologies & Methodologies
Education & Background
Academic and professional background
Specialization in Artificial Intelligence | Postgraduate Degree
Universidad Autónoma de Occidente
Advanced certification focused on machine learning and deep learning architectures, computer vision models, and production MLOps workflows. Core expertise includes deploying advanced algorithms, designing neural networks, and optimizing intelligent systems for real-world data processing applications.
IoT & AI for Healthcare | Research Internship
Universidad de Guadalajara
Designed and developed an Internet of Things (IoT) edge algorithm for real-time continuous monitoring of physiological variables under international scientific consultancy. embedded systems data pipeline efficiency, sensor integration, and secure data transmission protocols for critical healthcare applications.
Biomedical Engineering | Bachelor's Degree
Universidad Autónoma de Occidente
Undergraduate program focused on the intersection of healthcare technology, biological data science, and physical computing. Specialized in Computer Vision and medical image processing (DICOM/Radiomics), biomechanics keypoint tracking, and software development using machine learning and deep learning approaches.
Have an AI or Data Engineering project in mind?
Whether you need assistance building production-ready machine learning models, optimizing distributed ETL pipelines, or architecting cloud infrastructure—feel free to drop a message.