We combine local models and hosted ones to produce cheap tokens that rival the frontier.
imp is a new kind of inference provider. Our engine squeezes the most out of your personal machine to run capable local models efficiently, while occasionally falling back to hosted open models when local ones reach their limits.
imp comes as a lightweight daemon that plugs into your existing AI tools and applications, making them faster and absurdly cheaper to use without compromises. It runs on high-end machines equipped with Apple Silicon (with support for Nvidia DGX/RTX Spark coming soon), and selects optimal models based on your hardware.
We believe that balancing out on-device and hosted inference is an effective way to drive down the cost of intelligence and enable a new class of applications to be built. If this is something you'd like to work on, reach out.
imp is currently in development and will publicly launch later this year.