Juan('s re)search
Connecting Health, Genome and Data Science

Juan('s re)search
Connecting Health, Genome and Data Science
Juan is a computational scientist trained at MIT, UCSD and Weill Cornell Medicine with a PhD and 16 years of experience at the intersection of health science and data science, documented in over 67 peer reviewed publications with over 18 thousand citations.
The consistent theme in Juan's research (hence the name Juansearch) is identifying gaps in knowledge that can be bridged to meet the needs of underserved populations affected with diseases of large unmet need.
Hands-on advising of CEOs and PIs on strategy for implementation of AI an ML in health sciences. Prototypes, proof of concept studies, grant proposals, business plans, and manuscript revisions, along with Python and R notebooks for reproducibility.
Target discovery and genotype-first callbacks in founder population cohorts with exome sequencing and electronic medical records. Leading studies from concept to implementation, discovery to publication, target nomination to clinical validation.
Built extensive collaboration between Weill Cornell Medicine NY, Weil Cornell Medicine Qatar and Cornell University to advance the understanding of genetic disease risk in Arab populations. Led the development of QChip, first newborn screening array tailored to genetic diseases enriched in Arab populations.
Studied genetic basis of health disparities across genetically-defined human populations as part of 1000 Genomes Project Consortium.
Developed computational algorithms for classifying tumors based on gene expression profiles.
Funded by NIH/NHGRI F31 Kirschstein National Research Service Award. Thesis: "In silico, in vitro, in vivo and in populo: regulatory genetics of single nucleotide polymorphisms in the phenylethanolamine N-methyltransferase promoter." Awarded funding for travel to international workshops in Statistical Genetics and Gene Mapping.
Thesis: Identification of epistatic gene interactions by genome-wide double-knockout screening in Mycoplasma. Extensive coursework in theory of machine learning, artificial intelligence and statistics. TA for Databases and Algorithms course.
Developed experimental, imaging and analysis protocol for four-dimensional imaging of immunological synapse. Developed expertise in the logistics of government-funded research.
Developed websites and web applications on Linux/Apache/MySQL/PHP stack for startup companies and nonprofits.
Took leave of absence after 1st year.
National Honor Society. Presidential Academic Fitness Award. President of Natural Resources Society. Varsity Soccer and Swimming Team. 1st place, Science Bowl. Teaching Assistant for Tropical Ecology Summer Program.
Nearly decades connecting population genetics to disease mechanism, with an eye on translating these discoveries into diagnostic and therapeutic products. Published research includes work at Regeneron Genetics Center1-24 (RGC), Weill Cornell Medicine25-51 (WCM) and University of California at San Diego52-67 (UCSD).
At RGC, the focus of my work was building and expanding research collaborations in founder populations, including the Pakistan Genome Resource, Genome Ukraine, UK Biobank, Geisinger Health, Mexico City Prospective Study, and the Caribbean Genome Project. This work included discovery of novel therapeutic targets by ExWAS and genotype-first callbacks to predict the safety of existing targets.
At WCM, the focus of my research was a collaboration with WCM-Qatar, applying whole genome and exome sequencing to advance precision medicine through development of genomic medicine technology, discovery of novel disease-gene associations in multi-omic data, and in-depth exploration of the ancestry and population architecture of Arab populations. In parallel, I worked on 1000 Genomes Project collaborations using genomics to compare the ancestral mixture across Latin American populations.
At UCSD I conducted both wet-lab and computational research as part of my dissertation. The core of my work was development of computational models for predicting the impact of regulatory SNVs on gene expression and protein structure/function in the catecholamine synthesis pathway. My own research focused on PNMT, the adrenaline synthesis gene, where I showed the impact of promoter variants associated with hypertension and obesity on transcription factor binding and gene expression.
A body of work in human genetics, population genomics and translational biology, spanning first-author studies and large international collaborations.
Citation metrics and full publication history on Google Scholar.
U.S. Patent and Trademark Office, 2022 (US-20230000897-A1).
Selected publications
Hands-on support across the full arc of a research or translational project, from first question to publication-ready answer.
Population and statistical genetics, my core strength. Rare variants, founder populations, GWAS and EXWAS, admixture mapping, ancestry inference, multi-omics and functional genomics.
NGS workflows for germline and somatic variant discovery across WGS, WES, RNA-seq and epigenome. Built in Python, deployed to cloud or HPC.
Predictive models for OMICs, medical imaging and electronic health records, including LLMs and foundation models.
Study design, power calculations, mixed models and survival analysis in R or Python, with reproducible reports.
Target identification, biomarker discovery and translational research, from bench data to actionable insight.
New to pharma. I help entrants navigate first collaborations, scope joint projects and avoid the pitfalls that sink early deals.
Training the next generation of genomic scientists, from high school interns to faculty. Mentees have co-authored papers in Nature Communications and gone on to graduate programs and faculty appointments at Mt Sinai, Emory and Weill Cornell.
Whether you need a full-project collaborator or a single hour of strategic brainstorming, tell me about your challenge.