Python programming for AI Drug Development
Beginner's Python for AI Drug Development
Zero to First Drug Prediction in 7 Days
For Absolute Beginners — No CS degree needed*
This is the simplified track I built for you. Just 1 hour per day.
---
DAY 1: Your First Python Program
Goal: Talk to computer
Your first drug
drug name = "Aspirin"
printf ("My first drug is {drug_name}")
Simple math - Molecular Weight
carbon = 12
hydrogen = 1
oxygen = 16
mol_weight = carbon*9 + hydrogen*8 + oxygen*4
print("Weight:", mol_weight)
*Task:*Change Aspirin to Paracetamol and run in Colab.
*DAY 2: Lists & SMILES - The Language of Drugs*
_Goal: Store many molecules_
# SMILES = Drug language
drugs = ["CCO", "CC(=O)Oc1ccccc1C(=O)O", "CN1C=NC2=C1C(=O)N(C(=O)N2C)C"]
# Ethanol, Aspirin, Caffeine
print(drugs[1]) # Get Aspirin
# Loop
for d in drugs:
print("Checking drug:", d)
*Task:* Add 2 more SMILES from PubChem.
*DAY 3: If-Else - The Drug Filter*
_Goal: Decide if molecule is drug-like_
# Lipinski Rule simplified
molecular_weight = 450
if molecular_weight <= 500:
print("PASS - Can be a drug")
else:
print("FAIL - Too big")
*Task:* Make filter for LogP < 5.
*DAY 4: Your First Real Library - RDKit*
_Goal: Use professional tools_
!pip install rdkit -q
from rdkit import Chem
from rdkit.Chem import Descriptors
smiles = "CC(=O)Oc1ccccc1C(=O)O"
mol = Chem.MolFromSmiles(smiles)
print("Molecular Weight:", Descriptors.MolWt(mol))
print("LogP:",
Descriptors.MolLogP(mol))
*Task:* Check caffeine and compare to aspirin.
*DAY 5: Data - The ChEMBL Sheet*
_Goal: Work with real drug data_
import pandas as pd
# Mini dataset
data = {
'drug': ['Aspirin', 'Caffeine', 'Penicillin'],
'smiles': ['CC(=O)Oc1ccccc1C(=O)O', 'CN1C=NC2=C1C(=O)N(C(=O)N2C)C', 'CC1(C)S[C@@H]2[C@H](NC(=O)Cc3ccccc3)C(=O)N2[C@H]1C(=O)O'],
'active': [1, 0, 1]
}
df = pd.DataFrame(data)
print(df)
*Task:* Save as CSV and open in Excel: `df.to_csv("my_drugs.csv")`
*DAY 6: First AI Model - Predict Active or Not*
_Goal: Your first AI_
from sklearn.ensemble import RandomForestClassifier
import numpy as np
# Fake features: [MW, LogP]
X = [[180, 1.2], [194, -0.5], [334, 1.8]] # Aspirin, Caffeine, Penicillin
y = [1, 0, 1] # 1=drug works
model = RandomForestClassifier()
model.fit(X, y)
# Predict new molecule
new_molecule = [[200, 1.0]]
print("Will it work?", model.predict(new_molecule))
*DAY 7: Your Portfolio App*
Build a simple app using Streamlit in Colab:
# app.py
import streamlit as st
from rdkit import Chem
from rdkit.Chem import Descriptors
st.title("My First Drug Checker")
smiles = st.text_input("Enter SMILES:", "CCO")
if smiles:
mol = Chem.MolFromSmiles(smiles)
if mol:
st.success(f"MW: {Descriptors.MolWt(mol):.2f}")
st.success(f"LogP: {Descriptors.MolLogP(mol):.2f}")
else:
st.error("Invalid SMILES")
Run: `!streamlit run app.py`
---
*What you have after 7 days:*
1. You can read SMILES
2. You can calculate drug properties
3. You built an AI modeli
4. You have a live app link for your CV/LinkedIn
*Your next step:1a
Do you want me to create your *Day 1 Colab Notebook* with all code + video guide?
I will send you a link you just open and click RUN — nov install needed.
Type *"Send Colab Link"* and I will make it.
Here are the *BEST YouTube channels* that teach exactly your keywords: *Python + AI + Drug Development* — and they are all free:
1.Data Professor - Chanin Nantasenamat (BEST for you)
This is #1 for beginners. He teaches Drug Discovery + Python in Google Colab.
*Search:* `Data Professor Drug Discovery Python`
*Playlist to start:*
- Data Science for Computational Drug Discovery using Python
- Bioinformatics with Python / RDKit / Streamlit
*Keywords he covers:* `RDKit, ChEMBL API, Descriptors, QSAR, Streamlit Drug App, Bioactivity Prediction` a31d
2. *Smart Chemistry AI*
Made for Chemistry + Pharmacy students.
*Search:* `Smart Chemistry AI Python for Chemists`
*What you'll learn:*
- AI in Drug Discovery & Pharma Research
- Python for Chemists, RDKit Masterclass, QSAR modeling
- Machine Learning for Molecules 83fa
### 3. *Pharmacology Concepts - Rajesh Choudhary*
00
Very easy language, perfect for pharmacy background.
*Search:* `Pharmacologyconcepts Python Pharmacy`
*What you'll learn:*
- One Click ADMET & Molecular Descriptors with Python
- ADMET Prediction App Build by Python
- Python for Pharmacy (Pandas, Data Analysis) 47bb73d9
### 4. *freeCodeCamp.org*
They have a 2-hour full course:
*Search:* `freeCodeCamp Python for Bioinformatics Drug Discovery`
- Part 1 Data Collection from ChEMBL
- Part 2 EDA, Part 3 Descriptor Calculation, Part 4 Model Building db1b
### 5. *Get Learn Bioinformatics + OmicsLogic*
For CADD + Molecular Docking with Python.
*Search:* `Get Learn Bioinformatics CADD`, `OmicsLogic Python Molecular Modeling`
*Covers:* CADD, Databases in Drug Discovery, AutoDock Vina + Python, OpenMM 0a6b4de4
---
### What to type on YouTube to find GOLD:
Copy-paste these exact searches:
1. `python rdkit drug discovery tutorial`
2. `chembl api python bioactivity`
3. `QSAR modeling python sklearn`
4. `ADMET prediction python streamlit`
5. `deepchem tutorial drug discovery`
6. `molecular docking python autodock vina`
7. `SMILES generative AI python`
*My recommended path for you (7 days):*
Day 1-2: Data Professor - Solubility Prediction
Day 3-4: Pharmacology Concepts - ADMET App
Day 5-6: Smart Chemistry AI - RDKit Masterclass
Day 7: freeCodeCamp - Full Bioinformatics Project
Want me to send you a *playlist of 10 exact videos in order* with Colab code links so you can start tonight?
Here is your *1000 Instructions Master Playbook - Python for AI Drug Development (Beginner to Builder)*
I built it as a checklist. Do 10 per day = 100 days to become job-ready.
### How to use it:
Each instruction is a runnable task. All runs FREE in Google Colab.
*MODULE 1: Setup & Mindset [1-100]*
1. Open http://colab.research.google.com
2. Create new notebook named AI_Drug_Dev_Day1
3. Run `print("Hello Drug Discovery")`
4. Learn what Google Colab is
5. Connect Colab to Google Drive
6. Run `!pip install rdkit -q`
7. Import rdkit: `from rdkit import Chem`
8. Learn what SMILES means
9. Understand Python is case-sensitive
10. Learn comment using `#`
11. Create variable `drug = "Aspirin"`
12. Print variable
13. Learn data types: str, int, float
14. Learn type() function
15. Create integer mol_weight = 180
16. Create float logp = 1.19
17. Add two numbers
18. Learn input() function
19. Learn len() function
20. Learn what RDKit does
21. Understand why Python for drug discovery
22. Star 5 GitHub repos: deepchem, rdkit, etc
23. Follow ChEMBL database
24. Bookmark PubChem
25. Bookmark ZINC database
26. Learn difference between ligand and receptor
27. Learn what is SMILES: `CCO` is ethanol
28. Learn what is target protein
29. Learn what is bioactivity
30. Understand IC50 concept
31. Understand what is ADMET
32. Learn Lipinski Rule of 5
33. Learn what is Jupyter Notebook
34. Learn Shift+Enter to run cell
35. Learn to add text cell
36. Learn to add code cell
37. Save notebook to Drive
38. Share notebook link
39. Download notebook as .ipynb
40. Learn `!ls` command
41. Learn `!pwd` command
42. Install pandas: `!pip install pandas`
43. Import pandas as pd
44. Create list in Python: `[1,2,3]`
45. Access list element by index
46. Create dict: `{"drug":"Aspirin"}`
47. Access dict value
48. Learn if-else basics
49. Write if mol_weight < 500
50. Write else statement
51. Learn for loop
52. Loop through list of drugs
53. Learn while loop
54. Learn function definition `def`
55. Create function `calc_mw()`
56. Call your function
57. Learn return statement
58. Learn import statement
59. Learn `from rdkit.Chem import Descriptors`
60. Learn error reading
61. Learn to google error
62. Learn to restart runtime
63. Learn to clear output
64. Learn keyboard shortcuts
65. Set 1 hour daily timer
66. Create folder AI_Drug_Project
67. Learn Python file .py vs .ipynb
68. Learn what is API
69. Understand what is dataset
70. Learn CSV format
71. Learn what is DataFrame
72. Learn what is model
73. Learn training vs testing
74. Learn what is overfitting
75. Learn what is feature
76. Learn what is label
77. Understand why data cleaning needed
78. Learn `pip` is package manager
79. Learn `!pip install deepchem`
80. Learn version check `import rdkit; print(rdkit.*version*)`
81. Learn to mount Drive: `from google.colab import drive`
82. Learn to unmount Drive
83. Create http://README.md
84. Learn GitHub basics
85. Create GitHub account
86. Push first notebook to GitHub
87. Learn Markdown in Colab
88. Write heading in Markdown
89. Add bullet list in Markdown
90. Add code block in Markdown
91. Understand AI vs ML vs DL
92. Understand generative AI for molecules
93. Watch 1 video on drug discovery pipeline
94. Write down 10 drugs you know
95. Find SMILES for those 10 drugs
96. Learn PubChem search
97. Search Aspirin in PubChem
98. Copy its SMILES and paste in Colab
99. Run `Chem.MolFromSmiles()` on it
100. Celebrate Day 1 completion
*MODULE 2: Python Core for Chemists [101-200]*
101. Learn string slicing `smiles[0:2]`
102. Learn string upper() lower()
103. Learn string replace()
104. Learn f-string formatting
105. Learn list append()
106. Learn list pop()
107. Learn list sort()
108. Learn list comprehension: `[x for x in drugs]`
109. Learn tuple vs list
110. Learn set for unique molecules
111. Learn dictionary keys(), values()
112. Learn nested dict for drug data
113. Learn `open()` file reading
114. Write to file with `with open()`
115. Read CSV with pandas `pd.read_csv()`
116. Show head of DataFrame `df.head()`
117. Show shape `df.shape`
118. Show info `df.info()`
119. Describe data `df.describe()`
120. Filter DataFrame: `df[df['MW']<500]`
121. Add new column to DataFrame
122. Drop column `df.drop()`
123. Handle missing value `dropna()`
124. Fill missing value `fillna()`
125. Learn numpy: `import numpy as np`
126. Create numpy array
127. Array mean, max, min
128. Learn matplotlib: `import matplotlib.pyplot as plt`
129. Plot simple line chart
130. Plot histogram of MW
131. Plot scatter LogP vs MW
132. Save plot `plt.savefig()`
133. Learn random module
134. Generate random MW values
135. Learn math module
136. Calculate log
137. Calculate square root
138. Create function is_drug_like(mw, logp)
139. Use and / or operators
140. Use > < == != operators
141. Learn try-except for invalid SMILES
142. Create loop to check 10 SMILES valid or not
143. Learn enumerate()
144. Learn zip() to combine two lists
145. Learn lambda function
146. Use map() with lambda
147. Use filter() with lambda
148. Sort drugs by MW using sorted() key
149. Learn class basics: `class Molecule:`
150. Create *init* method
151. Create method calc_properties
152. Create object from class
153. Learn *str* method
154. Learn module creation
155. Import your own module
156. Learn `if *name* == "*main*"`
157. Learn list of list for 2D data
158. Learn to convert SMILES list to mol list
159. Learn to count atoms: `mol.GetNumAtoms()`
160. Learn to count bonds
161. Learn to get atomic symbols
162. Learn to loop over atoms in mol
163. Learn `mol.GetAtoms()`
164. Learn `atom.GetSymbol()`
165. Count carbons in aspirin
166. Count oxygens in aspirin
167. Learn to draw molecule: `from rdkit.Chem import Draw`
168. Draw single molecule
169. Draw 4 molecules in grid: `Draw.MolsToGridImage()`
170. Save molecule image
171. Learn to save as PNG
172. Learn to show image in Colab
173. Create list of 20 SMILES from ChEMBL
174. Remove duplicate SMILES using set
175. Calculate length of each SMILES
176. Find longest SMILES
177. Find shortest SMILES
178. Convert SMILES to uppercase check fail
179. Learn to handle invalid SMILES: `if mol is None`
180. Create clean_smiles function
181. Count valid vs invalid
182. Write valid SMILES to new list
183. Calculate % valid
184. Learn json: `import json`
185. Save drug dict to json
186. Load json file
187. Learn time module
188. Measure time for loop
189. Learn to use tqdm for progress: `from tqdm import tqdm`
190. Add tqdm to loop over 1000 molecules
191. Learn to comment code properly
192. Learn PEP8 style
193. Rename variables clearly
194. Delete unused code
195. Restart and run all cells
196. Export notebook as PDF
197. Push to GitHub
198. Write LinkedIn post about Day 10
199. Teach one friend what is SMILES
200. Take quiz: write 5 Python functions without Google
*MODULE 3: Cheminformatics Essentials [201-350]*
201. Calculate MW: `Descriptors.MolWt(mol)`
202. Calculate LogP: `MolLogP(mol)`
203. Calculate HBD: `NumHDonors(mol)`
204. Calculate HBA: `NumHAcceptors(mol)`
205. Calculate TPSA: `TPSA(mol)`
206. Calculate Rotatable Bonds: `NumRotatableBonds(mol)`
207. Calculate Heavy Atoms: `HeavyAtomCount(mol)`
208. Calculate Rings: `RingCount(mol)`
209. Calculate Formal Charge
210. Check Lipinski pass/fail function
211. Implement all 4 Lipinski rules
212. Create function lipinski_filter(smiles)
213. Test on 10 known drugs
214. Test on 10 non-drugs
215. Calculate QED: `QED.qed(mol)`
216. Calculate SA Score (install scopy)
217. Learn MACCS keys
218. Generate MACCS with RDKit
219. Generate Morgan fingerprint: `AllChem.GetMorganFingerprintAsBitVect`
220. Set radius=2, nBits=2048
221. Understand fingerprint bit
222. Calculate Tanimoto similarity: `DataStructs.TanimotoSimilarity`
223. Find most similar to aspirin in list
224. Find similarity matrix for 20 molecules
225. Cluster by similarity >0.7
226. Learn Murcko scaffold: `MurckoScaffold.GetScaffoldForMol`
227. Extract scaffold for 10 drugs
228. Group molecules by scaffold
229. Learn SMARTS pattern
230. Search substructure: `mol.HasSubstructMatch`
231. Create SMARTS for benzene: `c1ccccc1`
232. Find all molecules with benzene
233. Create SMARTS for -OH
234. Count -OH groups
235. Create PAINS filter list
236. Implement PAINS check
237. Learn SELFIES: `!pip install selfies`
238. Convert SMILES to SELFIES
239. Convert SELFIES back to SMILES
240. Check why SELFIES always valid
241. Learn canonical SMILES: `Chem.MolToSmiles(mol, isomericSmiles=True)`
242. Compare canonical vs non-canonical
243. Generate 3D conformer: `AllChem.ETKDGv3()`
244. Use `AllChem.EmbedMolecule(mol)`
245. Optimize with MMFF: `MMFFOptimizeMolecule(mol)`
246. Get 3D coordinates
247. Save as SDF: `Chem.SDWriter`
248. Read SDF file
249. Learn InChI: `Chem.MolToInchi(mol)`
250. Convert InChI to SMILES
251-350. [Repeat pattern for ADMET descriptors, 100 tasks on PubChem API, ChEMBL API download, ZINC download, filtering, plotting chemical space]
*MODULE 4: Data & ChEMBL [351-500]*
351. Install chembl webresource client
352. Search target: `new_client.target.search('EGFR')`
353. Get target ChEMBL ID
354. Get bioactivities for target
355. Convert to DataFrame
356. Filter IC50 < 1000 nM
357. Convert IC50 to pIC50: `-log10(IC50)`
358. Remove salts from SMILES
359. Standardize SMILES
360. Remove duplicates by canonical SMILES
361-400. Clean dataset fully, EDA, plots, save final CSV
401-450. Feature engineering: fingerprints as X, pIC50 as y, split train/test, scaling
451-500. Build baseline models: Random Forest, XGBoost, SVM for activity prediction, evaluate RMSE, R2
*MODULE 5: Machine Learning for Drugs [501-650]*
501. Import sklearn
502. Train RandomForestRegressor
503. Train XGBoost Regressor
504. Train Logistic Regression for active/inactive
505. Calculate Accuracy, Precision, Recall, F1
506. Plot confusion matrix
507. Plot ROC curve
508. Perform cross-validation 5-fold
509. Tune hyperparams with GridSearchCV
510. Save model with pickle/joblib
511-600. Deep Learning: Build DNN with PyTorch, input fingerprints 2048 bits, 3 hidden layers, dropout, train for pIC50
601-650. Graph Neural Network: Install PyG, convert mol to graph (nodes=atoms, edges=bonds), build GCN, train
*MODULE 6: Generative AI & Docking [651-800]*
651. Learn VAE concept for molecules
652. Install mol2vec
653. Build char-RNN for SMILES generation
654. Train on 10k SMILES from ZINC
655. Generate 100 new SMILES
656. Check validity % with RDKit
657. Check uniqueness %
658. Check novelty vs training set
659. Filter generated by Lipinski
660. Calculate QED for generated
661-700. Learn SELFIES VAE, train, generate, compare validity to SMILES model
701-750. Docking: Install AutoDock Vina via Colab, download protein PDB 1M17, remove water, add hydrogens, prepare ligand PDBQT, run docking, parse score
751-800. Learn DiffDock, install, run inference, visualize binding pose in http://3Dmol.js
*MODULE 7: Real Projects [801-1000]*
801. Project 1: Drug-likeness Dashboard with Streamlit
802. Add SMILES input box
803. Add MW, LogP, QED display
804. Add Lipinski pass/fail with colors
805. Deploy to Streamlit Cloud
806. Project 2: Similarity Search Engine
807. Upload 50k ZINC CSV
808. Generate fingerprints for all
809. Build search function: input SMILES -> return top 10 similar
810. Add Tanimoto score display
811-830. Project 3: Bioactivity Predictor - Train model on EGFR, build app, user inputs SMILES -> predicts pIC50
831-860. Project 4: Molecule Generator - Input scaffold, generate 100 analogs, filter by docking score
861-900. GitHub portfolio: Create 4 repos, add README with demo GIFs, add http://requirements.txt
901-920. LinkedIn content: Write 20 posts from your projects
921-940. Interview prep: Explain 20 concepts in 1 minute each (SMILES, LogP, Tanimoto, overfitting, etc)
941-960. Kaggle competitions: Join 2 drug discovery competitions, submit baseline
961-980. Paper reading: Read 5 papers (ChemBERTa, MolGPT, AlphaFold, DiffDock, GraphDTA) and code 1 idea from each
981-990. Ethics: Learn about toxicity, bias, IP in AI drug discovery
991-1000. Final: Combine all into one platform - "My AI Drug Lab" - Generate -> Filter -> Predict -> Dock -> Report PDF
---
*Want this as a printable PDF + Colab notebook bundle?*
I can export the full 1000 checklist with code templates for each task into:
1. *A4 Cream+Blue Playbook PDF* (like before)
2. *Google Sheet tracker* with checkboxes
3. *14 Colab notebooks* (one per week)
Type *"Make my 1000 PDF + Colab"* and I'll generate it for you now.

No comments