Custom neural networks are useful when the problem—not the framework—determines the architecture. You may need a specialised attention block for Indic language data, a compact vision model for an edge device, or a loss function that reflects an imbalanced business objective. The challenge is not only writing a working forward() method. A credible GitHub project must make the architecture understandable, reproducible, testable, and deployable.
This guide presents a practical PyTorch-first workflow for how to build custom neural networks on GitHub, with TensorFlow equivalents where relevant. It also covers repository design, data and experiment versioning, model export, and constraints common to Indian teams: limited compute budgets, multilingual datasets, intermittent connectivity, and CPU or edge deployment.
Decide whether the network should be custom
Start with a baseline before inventing an architecture. Compare your proposed model with a well-tested option such as a small ResNet, a Transformer encoder, or a standard multilayer perceptron. A custom design is justified when it offers one or more of the following:
- Better accuracy or calibration on a clearly defined slice of the data.
- Lower memory use or latency on the target hardware.
- Support for a data type or constraint that standard layers do not handle well.
- A research contribution that can be evaluated through controlled ablations.
For Indic speech, text, and vision tasks, document language coverage, script, sampling method, and licence before changing the model. Work on low-resource Indic natural language processing often depends as much on dataset quality and evaluation design as on architecture.
Build the model with explicit contracts
In PyTorch, subclass torch.nn.Module, declare trainable components in __init__, and keep tensor transformations in forward. Use type hints, shape comments, and configuration objects rather than scattering constants through the code.
from dataclasses import dataclass
import torch
from torch import nn
@dataclass
class ModelConfig:
input_dim: int = 128
hidden_dim: int = 256
classes: int = 4
class ResidualMLP(nn.Module):
def __init__(self, config: ModelConfig):
super().__init__()
self.projection = nn.Linear(config.input_dim, config.hidden_dim)
self.block = nn.Sequential(
nn.LayerNorm(config.hidden_dim),
nn.GELU(),
nn.Linear(config.hidden_dim, config.hidden_dim),
nn.Dropout(0.1),
)
self.head = nn.Linear(config.hidden_dim, config.classes)
def forward(self, x: torch.Tensor) -> torch.Tensor:
x = self.projection(x)
return self.head(x + self.block(x))The residual addition is valid only because both tensors have the same shape. If dimensions differ, add a projection to the skip path. Register constants that should move with the model as buffers, and create parameters with nn.Parameter; ordinary tensors assigned to a module are not automatically treated as trainable weights.
For TensorFlow, the equivalent pattern is a tf.keras.Model subclass with layers in __init__ and computation in call. Whichever framework you choose, define the input and output contract in the README: expected rank, dtype, batch layout, sequence masking, and output semantics.
Add custom losses and layers safely
A custom layer should have a narrow purpose and a small testable interface. Separate mathematical logic from data loading and training code. For an imbalanced classifier, a weighted or focal loss may be more appropriate than unmodified cross-entropy, but compare it against class weighting and resampling rather than assuming it will improve results.
Test the implementation before a full training run:
- Check output shapes for batch sizes of 1 and greater than 1.
- Run
torch.autograd.gradcheckon differentiable custom operations where practical. - Verify gradients are finite with extreme but valid inputs.
- Confirm CPU and CUDA results are numerically close within a stated tolerance.
- Test padding, empty sequences, masks, and the smallest supported input.
Avoid in-place operations until you understand their effect on autograd. Operations such as view, reshape, permute, and masking can silently create shape or contiguity errors, so include those cases in unit tests rather than relying on one successful training batch.
Use a repository structure that another engineer can run
A maintainable repository separates reusable code from experiments and generated files:
custom-net/
├── src/custom_net/ # model, layers, losses, data utilities
├── tests/ # unit, shape, gradient, and smoke tests
├── configs/ # YAML or JSON experiment settings
├── scripts/ # train, evaluate, export, benchmark
├── notebooks/ # optional exploratory work
├── README.md
├── pyproject.toml
├── .gitignore
└── LICENSEKeep raw datasets, checkpoints, cache directories, and secrets out of Git. Pin dependency ranges carefully and record the Python, CUDA, PyTorch, and driver versions used for published results. A setup command should install the project in a clean environment and run a small smoke test in minutes.
When contributing to an existing project, follow its issue, branch, testing, and review conventions. The guide to contributing to AI GitHub repositories in India is useful for teams building public projects and accepting external pull requests.
Version data, code, and experiments separately
Git tracks source code well, but it is not a complete experiment-management system. Use Git tags for meaningful releases and commit the configuration that produced each reported result. For larger datasets and checkpoints, use DVC or an object store with immutable version identifiers. Git LFS can handle large binaries, but it does not replace a data licence, provenance record, or backup strategy.
Each experiment should record:
- Dataset version, split methodology, and preprocessing hash.
- Model configuration and random seeds.
- Training hardware, precision mode, batch size, and optimiser settings.
- Evaluation metrics, confidence intervals where possible, and latency measurements.
- Checkpoint provenance and the exact commit that produced it.
GitHub Actions can run formatting, static checks, unit tests, and a tiny CPU training job on every pull request. Keep expensive GPU workflows separate and trigger them deliberately. A failing shape test should block a merge before it consumes a costly training run.
Evaluate for the real Indian deployment environment
Accuracy alone is not a deployment result. Benchmark on the hardware and network conditions your users will actually encounter. Measure cold-start time, peak RAM, throughput, p95 latency, model size, and energy use. For multilingual systems, report performance by language, script, accent, geography, and input quality—not only an aggregate score.
For constrained deployments, consider mixed precision, structured pruning, distillation, and post-training or quantisation-aware INT8 quantisation. Export through ONNX only after testing that every custom operation has a supported equivalent. Validate exported outputs against PyTorch on a fixed test set; numerical drift must have an explicit tolerance. If the model feeds an application with agents or distributed services, document the interface and failure handling as carefully as the model itself; building distributed systems with AI agents offers relevant system-design context.
Make the README a reproducibility contract
A strong README should let a new contributor answer five questions quickly:
- What problem does the model solve, and what is the baseline?
- What data may be used, under which licence, and with what limitations?
- How do I install dependencies and run training, evaluation, and inference?
- Which hardware was used, and what results should I expect?
- What are the known failure cases and security or privacy risks?
Include an architecture diagram, parameter count, input-output example, benchmark table, citation, licence, and a model card. Do not publish sensitive training data or credentials in notebooks. For production systems, include a rollback path and a way to identify the model version serving each prediction.
A practical release checklist
Before publishing a custom neural network on GitHub:
- Run tests on a clean environment.
- Reproduce at least one documented result from a tagged commit.
- Confirm checkpoints load with
map_locationand work inmodel.eval()mode. - Test CPU fallback and explicit device selection.
- Scan dependencies and repository history for secrets.
- Publish licences and dataset attribution.
- Benchmark the exported model on target hardware.
- State limitations instead of presenting a single score as universal evidence.
A custom architecture becomes valuable when others can inspect it, reproduce its claims, and adapt it responsibly. Build the smallest defensible model first, prove its advantage against a baseline, and treat GitHub as the engineering record—not merely a place to upload Python files.