AI Model Development

Explore top LinkedIn content from expert professionals.

  • View profile for Andrew Ng
    Andrew Ng Andrew Ng is an Influencer

    DeepLearning.AI, AI Fund and AI Aspire

    2,593,889 followers

    After a recent price reduction by OpenAI, GPT-4o tokens now cost $4 per million tokens (using a blended rate that assumes 80% input and 20% output tokens). GPT-4 cost $36 per million tokens at its initial release in March 2023. This price reduction over 17 months corresponds to about a 79% drop in price per year. As you can see, token prices are falling rapidly! One force that’s driving prices down is the release of open weights models such as Llama 3.1. If API providers, including startups Anyscale, Fireworks, Together AI, and some large cloud companies, do not have to worry about recouping the cost of developing a model, they can compete directly on price and a few other factors such as speed. Further, hardware innovations by companies such as Groq (a leading player in fast token generation), Samba Nova (which serves Llama 3.1 405B tokens at an impressive 114 tokens per second), and wafer-scale computation startup Cerebras (which just announced a new offering this week), as well as the semiconductor giants NVIDIA, AMD, Intel, and Qualcomm, will drive further price cuts. When building applications, I find it useful to design to where the technology is going rather than where it has been. Based on the technology roadmaps of multiple software and hardware companies — which include improved semiconductors, smaller models, and algorithmic innovation — I’m confident that token prices will continue to fall rapidly. This means that even if you build an agentic workload that isn’t entirely economical, falling token prices might make it economical at some point. Being able to process many tokens is particularly important for agentic workloads, which must call a model many times before generating a result. Further, even agentic workloads are already quite affordable for many applications. Let's say you build an application to assist a human worker, and it uses 100 tokens per second continuously: At $4/million tokens, you'd be spending only $1.44/hour – which is significantly lower than the minimum wage in the U.S. and many other countries. So how can AI companies prepare? - First, I continue to hear from teams that are surprised to find out how cheap LLM usage is when they actually work through cost calculations. For many applications, it isn’t worth too much effort to optimize the cost. So first and foremost, I advise teams to focus on building a useful application rather than on optimizing LLM costs. - Second, even if an application is marginally too expensive to run today, it may be worth deploying in anticipation of lower prices. - Finally, as new models get released, it might be worthwhile to periodically examine an application to decide whether to switch to a new model either from the same provider (such as switching from GPT-4 to GPT-4o-2024-08-06) or a different provider, to take advantage of falling prices and/or increased capabilities. [Reached lenght limit. Full text: https://lnkd.in/gz-xffF4 ]

  • View profile for Brij Kishore Pandey

    AI Architect & AI Engineer | Building Agentic Systems & Scalable AI Solutions

    736,795 followers

    The AI is coming back to your laptop. Google’s latest Gemma 4 12B launch is another strong signal of where AI infrastructure is heading: Local AI. For the last two years, most AI applications followed a simple pattern: User → App → Cloud API → Response But that pattern is changing. Now we are moving toward: User → Device → Local model → Cloud model when needed → Tools / Response This shift is important because not every AI task needs a frontier cloud model. Some tasks need speed. Some need privacy. Some need offline access. Some need lower cost. Some need to stay close to the user’s data. That is where local AI becomes powerful. Cloud AI is still critical for large-scale reasoning, heavy workloads, centralized governance, and enterprise reliability. Local AI is better for privacy-sensitive tasks, low-latency responses, personal productivity, offline workflows, and edge use cases. Hybrid AI is where most real-world systems will eventually land. The key architecture question will no longer be only: “Which model API should we use?” It will become: “Where should this intelligence run?” On the device? In the browser? In the private cloud? In the public cloud? Across all of them? This is a big mindset shift for AI engineers. Model selection is no longer enough. Model placement is becoming a core architecture decision. The winning systems will combine: small local models specialized edge models enterprise RAG systems cloud frontier models governed tool access human approval layers observability across the full workflow Local AI does not replace cloud AI. It changes where intelligence runs. And that means the laptop, phone, browser, and edge device are becoming part of the AI infrastructure stack. The next wave of AI engineering will be about routing the right task to the right model in the right place. The real question is: What should run locally, and what should go to the cloud?

  • View profile for Tobias Zwingmann
    Tobias Zwingmann Tobias Zwingmann is an Influencer

    Author of The Profitable AI Advantage. | Getting you to the edge of what AI can do without losing focus on what your business needs. | Instructor at LinkedIn Learning & O’Reilly Media

    85,348 followers

    Meta (Facebook) just spent over $100,000,000 to build Llama 3 - the world's best open source AI model - so you can use it for free. Here's why they did it and how you can benefit: Meta is the only Big Tech company committed to developing AI, particularly large language models, with an open-source approach. Why? My take: it's less about altruism and more about strategy. By making LLMs essentially free, they're aiming to crush all commercial competition. Llama 3, Meta's latest model, was trained on 24,000 (!) GPUs (each costing around $40k). You can do the math. Two versions are immediately available: - Small (GPT-3.5 class, 8B) - Medium (below GPT-4 class, 70B) The large model (GPT-4 class) is still in training and will be released later this year. I've played with the medium-sized model and I'm really impressed with the results. Meta has built a great open source model with really good reasoning and logic capabilities. There are 3 ways you can use Llama 3 for your business: 1/ Llama 3 as a Service Use Llama 3 from any cloud provider as a service. You pay by use, but the price is typically much cheaper than proprietary models like GPT-4 or Claude. → Use Llama 3 on Azure AI catalog: https://lnkd.in/eptszsUD 2/ Self-Hosting If you have GPU infrastructure (on-premises or cloud), you can run Llama 3 internally at your desired scale. → Deploy Llama 3 on Amazon SageMaker: https://lnkd.in/ext8cmBH 3/ Desktop (Offline):  Tools like Ollama allow you to run the small model offline on consumer hardware like current MacBooks. → Tutorial for Mac: https://lnkd.in/eKnPbHne Bottom Line: There’s really no point any more in trying to train your own large language model from scratch unless you’re in Big Tech or AI research. Just take the models that are out there, fine-tune on your data (if necessary) and start building! Will you use these models to augment workflows, increase productivity or simply be more creative in your business? The choice is yours - but the opportunity is too big to ignore. Have you tried Llama 3 yet?

  • View profile for Sol Rashidi, MBA
    Sol Rashidi, MBA Sol Rashidi, MBA is an Influencer
    120,669 followers

    Every board is betting big on AI. Almost none are asking the question that actually protects them. I’ve been in boardrooms across industries, from finance to healthcare, and I keep seeing the same thing: Board members ask: - “What’s the AI budget?” - “What’s the timeline?” - “What’s the ROI?” But almost no one asks the most important question: “How do we even know this is AI?” Here’s the problem… Most boards are approving AI initiatives without a clear definition of what qualifies as AI because the lines are blurry. Vendors show up with polished demos and pitch tools labeled “AI-powered.” But without clarity, boards end up greenlighting: ✗ Rule-based systems dressed up as intelligence ✗ Traditional software relabeled with buzzwords ✗ Proof-of-concept demos, not scalable AI infrastructure ✗ “AI-washed” features that don’t actually learn or adapt Before the next AI contract crosses your desk, ask leadership: → Where exactly does machine learning happen in this system? → How does it improve over time with use? → What data powers it, and who owns that data? → How much human intervention is required for results? Because the companies truly win with AI? They’re not the ones with the flashiest tools. They’re the ones whose boards can differentiate real intelligence from noise. What’s your take - have you seen “AI” claims fall apart under scrutiny?

  • View profile for Bertalan Meskó, MD, PhD
    Bertalan Meskó, MD, PhD Bertalan Meskó, MD, PhD is an Influencer

    The Medical Futurist, Global Keynote Speaker, Researcher and Author.

    372,315 followers

    BREAKING! The FDA just released this draft guidance, titled Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations, that aims to provide industry and FDA staff with a Total Product Life Cycle (TPLC) approach for developing, validating, and maintaining AI-enabled medical devices. The guidance is important even in its draft stage in providing more detailed, AI-specific instructions on what regulators expect in marketing submissions; and how developers can control AI bias. What’s new in it? 1) It requests clear explanations of how and why AI is used within the device. 2) It requires sponsors to provide adequate instructions, warnings, and limitations so that users understand the model’s outputs and scope (e.g., whether further tests or clinical judgment are needed). 3) Encourages sponsors to follow standard risk-management procedures; and stresses that misunderstanding or incorrect interpretation of the AI’s output is a major risk factor. 4) Recommends analyzing performance across subgroups to detect potential AI bias (e.g., different performance in underrepresented demographics). 5) Recommends robust testing (e.g., sensitivity, specificity, AUC, PPV/NPV) on datasets that match the intended clinical conditions. 6) Recognizes that AI performance may drift (e.g., as clinical practice changes), therefore sponsors are advised to maintain ongoing monitoring, identify performance deterioration, and enact timely mitigations. 7) Discusses AI-specific security threats (e.g., data poisoning, model inversion/stealing, adversarial inputs) and encourages sponsors to adopt threat modeling and testing (fuzz testing, penetration testing). 8) And proposed for public-facing FDA summaries (e.g., 510(k) Summaries, De Novo decision summaries) to foster user trust and better understanding of the model’s capabilities and limits.

  • View profile for Maximus Friedrich Baluyot

    AI & Automation Architect | Multi-Agent Systems | CRM Pipelines | Caltech Certified | 12+ Years

    3,634 followers

    Ever wondered how far you can push real-time computer vision using just a lightweight language model and your browser? I just explored smolvlm-realtime-webcam — a fascinating project that captures webcam input, sends it to a local llama.cpp server running SmolVLM (500M params), and gets back live object descriptions from a tiny vision-language model — all in real time. This isn't your typical deep-learning pipeline. It's: Extremely lightweight — no massive GPU needed Browser-based — just HTML + JS Powered by llama.cpp — fast inference on CPU/GPU Hackable — you can prompt it to return structured data like JSON Perfect for edge computing, fast prototyping, or simply geeking out on vision+language systems with minimal overhead. Big shoutout to @ngxson (Xuan-Son Nguyen) and the open-source community behind this. Want to see a llama do object detection from your webcam? Check it out: https://lnkd.in/gHB62wzY #AI #ComputerVision #EdgeAI #llama #SmolVLM #OpenSource #RealTimeAI #llamacpp #MachineLearning #TechDemo

  • View profile for Greg Coquillo

    AI Platform & Infrastructure Product Leader | Scaling massive AI Factories for Frontier Model providers | Azure AI & HPC | Former AWS, Amazon | Startup Investor | I deploy GPU-as-a-Service for AI customers

    234,344 followers

    AI-assisted coding isn’t just about autocomplete anymore. It’s becoming a full lifecycle - from planning to building to reviewing. Developers are no longer just writing code, they’re orchestrating systems of agents that generate, test, and refine it. The shift is from “write code faster” to “build and ship systems end-to-end.” Here’s how the generative programmer stack is evolving 👇 𝗕𝗨𝗜𝗟𝗗 - 𝗖𝗼𝗱𝗲 𝗚𝗲𝗻𝗲𝗿𝗮𝘁𝗶𝗼𝗻 & 𝗘𝘅𝗲𝗰𝘂𝘁𝗶𝗼𝗻 Full-Stack App Builders: Turn ideas into working applications quickly by generating frontend, backend, and integrations in one flow. CLI-Native Agents: Work directly from the terminal to generate, edit, and execute code with tight control and speed. IDE-Native Agents: Integrate inside development environments to assist with coding, debugging, and real-time suggestions. Async Cloud Coding Agents: Run tasks in the background - writing, testing, and iterating on code without blocking your workflow. 𝗣𝗟𝗔𝗡 - 𝗣𝗹𝗮𝗻𝗻𝗶𝗻𝗴 & 𝗙𝗲𝗮𝘁𝘂𝗿𝗲 𝗕𝘂𝗶𝗹𝗱𝗶𝗻𝗴 Spec-first Tools: Start with structured specifications that define what to build before writing any code. Ask / Plan Modes: Break down problems, explore approaches, and validate logic before jumping into implementation. Design-to-Code Inputs: Convert designs or structured inputs into working code, reducing manual translation effort. 𝗥𝗘𝗩𝗜𝗘𝗪 - 𝗥𝗲𝘃𝗶𝗲𝘄, 𝗧𝗲𝘀𝘁𝗶𝗻𝗴 & 𝗩𝗲𝗿𝗶𝗳𝗶𝗰𝗮𝘁𝗶𝗼𝗻 Code Review Agents: Automatically analyze code for issues, improvements, and best practices before deployment. Testing & Verification: Generate and run tests to ensure reliability, correctness, and stability across different scenarios. Benchmarks: Measure performance and quality using standardized evaluation frameworks. What this means: Coding is shifting from manual effort to guided execution. The developer’s role is moving toward direction, validation, and system design. The edge is no longer just writing better code. It’s knowing how to use these tools together to ship faster and more reliably. Which part of this workflow are you using AI for the most today?

  • View profile for Homayoun Rezaie

    AI 4 EARTH 🛰️ | PhD Candidate @Ucalgary

    17,114 followers

    NVIDIA just open-sourced a whole family of weather and climate ai models, and I think this is the moment serious #forecasting stops being something only national weather agencies can do. #Earth2 isn't one model, it's a stack. #Atlas does 15-day global forecasts and beats GenCast on benchmarks. #StormScope is the first AI to outperform physics-based systems on storm dynamics. #HealDA spins up initial atmospheric conditions in seconds on a GPU, the kind of step that used to take hours on a supercomputer. #CorrDiff downscales 500x faster while using 10,000x less energy. what gets me excited isn't any one of these though, it's that the whole pipeline is just sitting on #HuggingFace and #GitHub now. running weather ai used to mean physics models, now a small team in any country can fine-tune these and run them on their own hardware. bigger models are interesting, but this feels more important to me, taking something that lived behind a national lab's firewall and just handing it out. Link to Earth-2: https://lnkd.in/exReKuTV #climateAI #remotesensing #foundationmodel #weather #climate

  • View profile for Rahul Agarwal

    Staff ML Engineer | Meta, Roku, Walmart | 1:1 @ topmate.io/MLwhiz

    46,132 followers

    Few Lessons from Deploying and Using LLMs in Production Deploying LLMs can feel like hiring a hyperactive genius intern—they dazzle users while potentially draining your API budget. Here are some insights I’ve gathered: 1. “Cheap” is a Lie You Tell Yourself: Cloud costs per call may seem low, but the overall expense of an LLM-based system can skyrocket. Fixes: - Cache repetitive queries: Users ask the same thing at least 100x/day - Gatekeep: Use cheap classifiers (BERT) to filter “easy” requests. Let LLMs handle only the complex 10% and your current systems handle the remaining 90%. - Quantize your models: Shrink LLMs to run on cheaper hardware without massive accuracy drops - Asynchronously build your caches — Pre-generate common responses before they’re requested or gracefully fail the first time a query comes and cache for the next time. 2. Guard Against Model Hallucinations: Sometimes, models express answers with such confidence that distinguishing fact from fiction becomes challenging, even for human reviewers. Fixes: - Use RAG - Just a fancy way of saying to provide your model the knowledge it requires in the prompt itself by querying some database based on semantic matches with the query. - Guardrails: Validate outputs using regex or cross-encoders to establish a clear decision boundary between the query and the LLM’s response. 3. The best LLM is often a discriminative model: You don’t always need a full LLM. Consider knowledge distillation: use a large LLM to label your data and then train a smaller, discriminative model that performs similarly at a much lower cost. 4. It's not about the model, it is about the data on which it is trained: A smaller LLM might struggle with specialized domain data—that’s normal. Fine-tune your model on your specific data set by starting with parameter-efficient methods (like LoRA or Adapters) and using synthetic data generation to bootstrap training. 5. Prompts are the new Features: Prompts are the new features in your system. Version them, run A/B tests, and continuously refine using online experiments. Consider bandit algorithms to automatically promote the best-performing variants. What do you think? Have I missed anything? I’d love to hear your “I survived LLM prod” stories in the comments!

  • View profile for Montgomery Singman
    Montgomery Singman Montgomery Singman is an Influencer

    Managing Partner @ Radiance Strategic Solutions | xSony, xElectronic Arts, xCapcom, xAtari

    28,015 followers

    A team of researchers from Google Research, Google DeepMind, and Tel Aviv University has developed a groundbreaking AI application capable of recreating and simulating parts of existing video games, including the iconic game Doom. In a fascinating advancement for gaming and AI, researchers have modified a machine learning model to recreate video game environments and actions. Named GameNGen, this new system uses neural rendering techniques based on diffusion models to simulate realistic gameplay. The team trained the AI by feeding it video footage of Doom, allowing it to generate new gameplay frames nearly indistinguishable from the original. This development marks a significant step in the intersection of AI and gaming, opening up new game development and simulation possibilities. 🎮 Recreating Games with AI: The research team successfully used a modified diffusion model, GameNGen, to simulate sections of the video game Doom, highlighting the potential of AI in game development. 🧠 Neural Rendering Techniques: The process relies on neural rendering, where AI learns to recreate the imagery and the actions within a game, pushing the boundaries of what AI can achieve. 🖼️ Diffusion Models in Action: Building on the Stable Diffusion 1.4 model, GameNGen is explicitly trained on video game footage, allowing it to generate new, realistic gameplay frames. ⚙️ Realistic Gameplay Simulation: The AI-generated frames were shown to human raters, who often could not distinguish them from real game footage, demonstrating the model's effectiveness. 🚀 Impact on Game Development: This technology could revolutionize the gaming industry by enabling more efficient game development and even the possibility of creating entirely new games through AI. #GameNGen #AIinGaming #NeuralRendering #MachineLearning #VideoGameAI #GenerativeAI #DoomSimulation #GoogleResearch #DeepMind #GameDevelopment 

Explore categories